Proposal writing
Horizon Europe evaluation criteria, with writing examples for each section
12 min read · Updated 28 August 2026
Collaborative Horizon Europe proposals are scored against three criteria — Excellence, Impact, and Implementation — each out of 5, with half-point increments and thresholds. Understanding what sits behind each score, and writing to it explicitly, is worth more than another round of polishing the science.
How scoring actually works
- Each criterion is scored 0–5 in half points, normally with a threshold of 3 per criterion and an overall threshold of 10.
- Several evaluators score independently, then reach a consensus score and write a consensus report.
- In many Pillar II topics Impact carries extra weight in ranking, and Impact is where most proposals lose ground.
- In practice, the funding line sits well above threshold. A 12/15 proposal is a good proposal that is not funded.
Excellence — what is assessed
Excellence covers objectives, the relationship to the state of the art, the soundness and ambition of the methodology, and how the concept goes beyond current practice. Evaluators penalise vagueness and unsupported novelty claims.
- Objectives that are specific, measurable, and mapped to work packages.
- An explicit state-of-the-art comparison naming competing approaches.
- A methodology with justified choices, not a list of activities.
- Interdisciplinarity, open science practices, and gender dimension addressed where relevant to the content.
Excellence — writing examples
- Weak: "The project will develop an innovative AI-based platform that significantly improves current monitoring approaches."
- Stronger: "O1: reduce false-positive rate in continuous grid-fault detection from the 12–15% reported by [state-of-the-art method] to below 4%, validated on three operator datasets covering 18 months of operation (WP3, M18)."
- Weak: "There is no comparable solution on the market."
- Stronger: "The closest approaches are X (rule-based, needs manual recalibration per site) and Y (learning-based, requires labelled fault data unavailable in most networks). Our approach removes both constraints by [mechanism], which has not been demonstrated beyond laboratory scale."
- Weak: "We will use machine learning to analyse the data."
- Stronger: "We use a self-supervised model because labelled fault events are rare (< 0.1% of samples); the alternative, supervised classification, was ruled out in our preliminary study (Section 1.3) where it failed below 5,000 labelled events."
Impact — what is assessed
Impact covers the credibility of the pathway from your results to the expected outcomes and wider impacts in the topic text, the scale and significance of the contribution, and the measures for dissemination, exploitation, and communication, including IP management.
- A pathway with logical steps: results → outcomes → impacts, each with a named beneficiary.
- Quantified targets with a baseline, a method of measurement, and a timeframe.
- Barriers to uptake named honestly — regulation, cost, procurement cycles, standards — with a response.
- Concrete key exploitable results with owners, routes to market, and IP arrangements.
Impact — writing examples
- Weak: "The project will contribute to the European Green Deal and improve energy efficiency across Europe."
- Stronger: "Outcome: distribution system operators cut unplanned outage minutes. Baseline: 78 min/customer/year across the three pilot operators (2025 regulator data). Target: −15% within two years of deployment, measured from the same regulatory reporting. Reached via the operator association's 41 members (letter of support, Annex)."
- Weak: "Results will be disseminated widely through conferences, publications and social media."
- Stronger: "KER2 (fault-detection engine, owned by Partner B, background IP declared in the CA) reaches market through Partner B's existing SCADA integration channel; 4 operator workshops (M18, M24, M30, M36, 25 participants each) target procurement decision-makers, not researchers."
- Weak: "The solution will be scalable to other sectors."
- Stronger: "Transfer to water networks requires only replacing the sensor adapter layer; this is assessed in T6.4 with one water utility as an associated partner, keeping the claim testable within the project."
Implementation — what is assessed
Implementation covers the quality and effectiveness of the work plan, the allocation of resources, and the fit of the consortium. Evaluators are looking for a plan that a competent coordinator could actually run.
- A work package structure with clear dependencies and no orphan tasks.
- Milestones that are decision points, and deliverables that are outputs — not the same list twice.
- A risk register with likelihood, impact, and a specific mitigation owner.
- Effort and budget that match the described work, with no partner present only for geographic balance.
- Management structures proportionate to consortium size, including decision-making and conflict resolution.
Implementation — writing examples
- Weak: "Risk: delays in data collection. Mitigation: the consortium will closely monitor progress."
- Stronger: "Risk R4: operator data-sharing agreements not signed by M6 (likelihood medium, impact high — blocks T3.2). Mitigation: two operators already signed (Annex); a synthetic dataset generated in T3.1 allows model development to continue for up to 5 months. Owner: WP3 leader."
- Weak: "Partner E will support dissemination activities."
- Stronger: "Partner E (sector association, 41 member operators) leads T7.2 and hosts the 4 procurement workshops; 6 PM allocated, matching its role in comparable projects [ref]."
- Weak: "MS3: WP4 completed."
- Stronger: "MS3 (M24): pilot achieves < 6% false positives on operator data; if not met, the consortium triggers the fallback architecture defined in T4.5 rather than entering the second pilot phase."
A final pass before submission
- Read the topic text again and highlight every expected outcome. Each one should appear, in the topic's own words, in your Impact section.
- Check that every claim with a number has a source, a baseline, or a method behind it.
- Have someone outside the field read the first two pages and explain the project back to you.
- Check consistency: effort tables, budget, Gantt, and narrative must agree.
- Leave the last week for consistency, not for new content.
Frequently asked questions
- What score do I need to be funded?
- Thresholds are usually 3 per criterion and 10 overall, but in competitive topics the funded range typically starts around 13.5–14.5 out of 15. Aim for the top of the range, not the threshold.
- Which criterion should I spend the most time on?
- Impact, in most collaborative calls. Research teams are usually strongest on Excellence and weakest on the pathway from results to measurable outcomes — and Impact is often decisive in ranking.
- Do evaluators read the annexes?
- They read what the page-limited part directs them to. Never place an essential argument only in an annex; use annexes for evidence that supports a claim already made in the main text.
- How closely do I need to match the topic wording?
- Closely. Evaluators score against the topic's expected outcomes and scope. Using your own vocabulary for the same concepts makes the match harder to see and costs points that are entirely avoidable.