SEO
GEO agencies in France: only 9 of 30 publish verifiable evidence
Edikka Observatory · 30 websites · 18 checks · Aggregate data
France’s GEO market knows how to frame a promise. This study measures what a prospective client can actually verify, without publishing a ranking or any individual result.
- Sample 30 websites selected before measurement through five frozen queries.
- Cases 16 websites describe a GEO case or experiment.
- Evidence 9 out of 30 publish something a third party can verify.
- Protection No name, domain, score or individual remediation advice is published.
Two technical passes conducted more than 18 hours apart, 300 documented editorial decisions and a public file limited to aggregates from the 18 checks.
Short answer
Only 9 out of 30 websites publish GEO evidence that a third party can verify.
The 30 websites in the sample qualified because they presented a GEO or AI search optimization service. Their ability to name the market is therefore an eligibility criterion, not a finding. The real insight emerges further down the funnel, when we look for a structured method, a case explicitly connected to GEO, a dated measurement, a baseline and something an independent party can inspect.
Edikka observed the websites of 30 providers selected through five search queries fixed before data collection began. There is no league table, no overall score out of 100, no published names and no judgement about capabilities that are not publicly documented: just eight technical checks and ten binary editorial checks, reported in aggregate.
The central finding does not answer “which is the best GEO agency?”. It answers a more useful question: what can a prospective client actually verify before choosing an agency?
16 websites describe a GEO case or experiment, but only 9 publish a named client, queries, a dataset, a contextualized screenshot or another independently verifiable element.
Pre-registered method
The sample was selected before scoring, using a deterministic cut across five queries.
Five French-language search formulations were frozen before collection: “agence GEO” France, “agence référencement IA”, “agence SEO IA”, “audit GEO” référencement and “agence Generative Engine Optimization”. The first 12 web results for each query were recorded with their rank, URL and timestamp.
The sample was then assembled by taking one result from each query in turn. The same commercial entity could appear only once, even if it operated several domains. Standalone SaaS products, media outlets, directories, rankings without an associated service, and solo consultants for whom no business structure could be observed were excluded; the reason for each exclusion was retained in the snapshot.
The full method, aggregates, denominators, protocol fingerprints and public changelog accompany the study. Individual results remain internal: the measurement is transparent, without turning the Observatory into a ranking of competitors.
| Dimension | Question | Checks | What this dimension does not prove |
|---|---|---|---|
| Public foundation | Are the website and its service technically readable? | HTTP, indexability, canonical, sitemap, internal link, crawlers, Schema.org and multi-page coverage. | Neither the quality of the advice nor client outcomes. |
| Public evidence | What can a prospective client verify before a sales call? | Service, deliverables, method, metrics, limitations, case, figure, duration, baseline and verifiability. | Neither private methods nor unpublished client satisfaction. |
A single score would allow a good sitemap to compensate artificially for the absence of a public case. The two dimensions therefore remain separate: a technical foundation scored out of 8 and editorial evidence scored out of 10, with no podium.
Results
The drop-off between a service page and verifiable evidence is unmistakable.
Among decisive results, all 30 websites present an identifiable GEO service and 29 out of 30 name at least two AI visibility metrics. The market has broadly learned how to talk about the subject.
27 publish a method with at least three stages. 21 explain at least one limitation: variable answers, dependence on the engine, no guarantee or partial measurement. 16 describe a case or experiment, 13 document a baseline or protocol, and 9 let a third party verify at least one element.
| Check | Observed | Decisive denominator | Interpretation |
|---|---|---|---|
| Public GEO service | 30 | 30 | Commercial positioning is no longer rare. |
| Structured method | 27 | 30 | Methods are usually published; whether they are verifiable is a separate question. |
| Stated limitation | 21 | 30 | Uncertainty is still disclosed less often than the promise. |
| Case or experiment | 16 | 30 | Evidence exists, but its depth varies considerably. |
| Baseline or protocol | 13 | 30 | Without a starting point, improvement remains difficult to interpret. |
| Verifiable evidence | 9 | 30 | This is the main bottleneck. |
What really matters
A marketing figure is not yet GEO evidence.
An increase in traffic, a client count or a promise about timing is not enough if the result is not connected to visibility in AI-generated answers. Conversely, a modest experiment can be highly valuable when it specifies the prompts, engines, dates, starting point and observed outputs.
The benchmark also accepts self-experiments and aggregate analyses. It does not claim they all carry equal weight: that is precisely why “published case”, “quantified result”, “baseline” and “verifiable evidence” are four separate checks.
Context
Name the website, industry or object of the test.
A reader should know what was optimized and under what conditions.
Protocol
Publish the prompts, engines, dates and repetitions.
A single answer does not measure probabilistic visibility.
Comparison
Retain a before state, baseline or reference group.
A result without an initial state cannot establish that progress occurred.
Verification
Leave a trace that a third party can inspect.
A named client, dataset, dated screenshot, published queries, report or rerun URL.
Decision guide
Five questions separate a GEO promise from evidence you can use.
The aim is not to choose a provider from a table. It is to improve the questions asked before making a commitment. A serious answer should connect an action, a measurement and a verifiable trace.
| Question | What to expect | Insufficient on its own |
|---|---|---|
| What exactly do you measure? | Citations, mentions, cited pages, share of voice or AI referral traffic. | “We improve your AI visibility.” |
| Across which corpus? | Prompts, engines, languages, markets and observation frequency. | One isolated query in ChatGPT. |
| What is the starting point? | A dated baseline or before-and-after comparison. | A final result with no initial state. |
| Which limitations do you disclose? | Variability, covered engines, timing, context and the absence of guarantees. | A promise of a stable position. |
| What can a third party verify? | A named client, queries, contextualized screenshot, report or dataset. | A percentage with no source or protocol. |
Wave 1 is designed to make the market easier to understand without handing out scores, diagnoses or remediation advice to competitors. The public file contains one row per check and no individual results.
Technical foundation
An llms.txt file carries no weight in this benchmark.
The scanner records crawl policies, sitemaps, canonicals and structured data, but awards no point for llms.txt. Google states in its documentation on AI features that no special markup or new machine-readable file is required: technical and editorial fundamentals remain the foundation.
Giving llms.txt no weight also prevents Edikka from benefiting from a signal it has documented extensively itself. A website can be highly readable without llms.txt; an llms.txt file cannot compensate for an undiscoverable service, a noindex directive or missing evidence.
To assess your own website, use the AI agent readiness test. For the trade-offs involved in crawler access, read the guide to blocking AI crawlers.
OpenAI recommends allowing OAI-SearchBot so that a website can be discovered and cited in ChatGPT Search. This access condition is nevertheless distinct from evidence of GEO effectiveness.
Microsoft now distinguishes citations, cited pages and grounding queries in Bing Webmaster Tools. That separation reinforces a central idea of the Observatory: being technically accessible, being cited and being able to interpret that citation are three different questions.
Limitations and governance
This benchmark measures what is public at a given date, not the full quality of an agency.
An agency may have robust methods, confidential outcomes or documents shared only with prospects. These are not counted because a third party cannot inspect them under the same conditions. Conversely, a published claim is recorded without Edikka automatically certifying its commercial accuracy.
The two technical scans were conducted at least 18 hours apart. A discrepancy remains “inconclusive”. Evidence, URLs and individual decisions are retained in the internal ledger, but wave 1 publishes no name, domain, score or third-party excerpt.
The sample reflects one web-search surface and five query formulations, not the entire French market. The next wave will revisit the same baseline after 90 days; any change to the method will be versioned separately.
The study says “not publicly observed”, never “the agency cannot do it”.
Conclusion
The GEO market does not need another ranking. It needs an evidence standard.
Across the observed sample, the vocabulary, services and new KPIs have already been adopted widely. The next step is to make results easier to interpret through context, protocols, dates, limitations and verifiable elements.
That is the role of this Observatory: not to declare a winner, but to make the market more legible. In 90 days, wave 2 will be able to measure changes in aggregate rates without exposing any competitor’s individual roadmap.
Edikka applies this standard to its own engagements: never sell a promise of being cited, but a protocol, dated measurements and decisions the client can verify.
Edikka’s perspective
GEO will mature when its evidence becomes easier to verify than its promises.
The Observatory does not choose an agency for you. It gives decision-makers a shared language for examining method, measurement, limitations and traceability.
Use binary checks
Public rules, never a taste-based score awarded by a competitor.
Retain sources
Every observed result retains a URL, excerpt, date and validation.
Run it again
The next wave measures change with a versioned method, not a frozen league table.
Go further on this topic
Additional answers to clarify the key points covered in this article.