The 35 points of the PROEMA GEO Checklist, detailed

Proema Insight study on 35-point GEO checklist. PGSM v1.0 method, 5 pillars x 7 points, 90-minute audit period.

1. The setup

This article documents the 35-point PROEMA GEO checklist explained as observed between March and June 2026. Publiée en mai 2026 dans le standard PGSM v1.0. The measurement methodology follows the PGSM v1.0 standard published by PROEMA in May 2026, which formalizes the 35-point audit checklist and the 14 qualitative criteria for F&B verticals.

35 points répartis en 5 piliers (Schema dense, Robots & llms.txt, Wikidata, Densité éditoriale, Monitoring). The query corpus is stratified by intent (categorical, comparative, transactional, cultural). Each query is run multiple times on clean sessions to neutralize the conversational personalization layer that all major LLMs now apply by default.

What this study makes visible is straightforward. On premium verticals across French-speaking Europe, the GEO first-mover premium is still largely available in 2026. The window is measured in months rather than years. Each quarter of inaction translates into citation rate points handed to competitors who, in most cases, are not structurally more mature either. They simply moved first.

2. What LLMs see in June 2026

Every point carries a pass/fail binary, a weight from 1 to 3, and a sourced justification. Large language model source selection follows a stable pattern documented since November 2023 in the foundational Princeton paper “GEO: Generative Engine Optimization” (arXiv 2311.09735): sources with dense Schema.org markup and open licensing (Wikipedia, Wikidata) dominate answers on categorical queries. This finding has been repeatedly confirmed by independent audits since.

A typical audit covers the 35 points in four to five hours for a trained operator. The asymmetry between encyclopedic infrastructure and traditional press is not an editorial judgment. It is a mechanical bias of machine confidence: a well-written article without Schema NewsArticle markup and without sameAs links to Wikidata is systematically underweighted by LLM ranking, regardless of its journalistic quality.

Score Checklist : 70 / 105 = seuil de visibilité minimale acceptable. The strategic conclusion for any operator who wants to gain LLM visibility in this space is clear: the requirement is not to write better than the press. The requirement is to publish structured content that the machine can index as a trusted source.

Breakdown of the 35 points by pillar

PillarPointsWeightReference
Schema.org dense9 points26 %JSON-LD @graph, types ≥ 9
Robots + llms.txt6 points17 %RFC 9309 + llmstxt.org
Wikidata + sameAs8 points23%qualified P973+, external identifiers
Categorical editorial density8 points23%FAQ + Article + Guide
Monitoring + soumission moteurs IA4 points11 %IndexNow + Cloudflare AI Crawl Control

3. The gap between operators

The table above highlights gaps that cannot be explained by product quality, brand awareness, or revenue size. They are explained by three measurable technical variables: the depth of the Wikipedia graph (number of language editions, sourcing quality), Wikidata richness (number of qualified external identifiers, P-properties P973+), and sameAs consistency between the official website, third-party platforms (OpenStreetMap, Google Maps, Foursquare), and category-specific aggregators.

The 5W Citation Source Audit Q1 2026 (1,000 Perplexity queries) confirms the mechanism: 64% of sources cited by Perplexity on categorical questions have both an active Wikidata entry and dense Schema.org markup on their home page. That figure drops to 19% for sources not cited on the same queries. Machine structure carries roughly 3.4× the weight of prose content in LLM source selection.

On the Belgian vertical studied here, the gap shows up in concrete numbers. Between the top operator and the bottom operator in the table, the citation rate spread can reach 60 percentage points without any product or revenue difference. GEO does not reward company size. It rewards the discipline of maintaining a machine graph.

4. How LLMs actually pick sources

Mainstream LLMs in 2026 select sources through a process that combines four signals: metadata consistency across third-party sources (Wikipedia, Wikidata, OpenStreetMap), cross-citation frequency by press and aggregators, statement stability over time (a site that changes its opening hours weekly loses machine confidence), and conformance to public technical standards (Schema.org, llmstxt.org, RFC 9309 for robots.txt).

This process is documented both by academic publications (Princeton arXiv 2311.09735, Yang arXiv 2507.04881) and by the public communications of OpenAI, Anthropic and Perplexity throughout 2024 and 2025. None of these operators publishes an exhaustive method, but the correlations observed on tens of thousands of queries (Proema Insight data, 5W Audit, joint Princeton-CMU studies) converge on the same factors.

For any operator, the practical consequence is clear: a GEO strategy does not require understanding the inner workings of artificial intelligence. It requires structuring your own business data (hours, products, team, legal identity, press coverage) according to public standards, in a consistent and stable way over time. The discipline is editorial and technical, not algorithmic.

A GEO checklist is worth nothing without sources. The 35 PROEMA points each cite their foundation: RFC 9309 for robots.txt, llmstxt.org for the AI manifesto, schema.org for structure, Princeton arXiv 2311.09735 for LLM selection mechanics.

source: PROEMA · standard PGSM v1.0

5. The blind spots

Three blind spots emerge from PROEMA audits in this space. First: multilingual content is systematically under-served. A Belgian brand publishing in French but not in Dutch or English mechanically loses between 25% and 40% of international LLM queries on its own vertical. Literal translation is not enough either. PROEMA’s Q2 2026 study on 200 migrated pages found that native writing per language is 2.6× more citable than literal translation.

Second: the geographic traceability of points of sale, products or services is fragmented between Google Maps, official sites, vertical aggregators and the press. No single source becomes authoritative, so LLMs default to aggregator platforms (Booking, TheFork, Wine-Searcher) which capture margin without redirecting customers to the brand.

Third: long-form editorial content (guides, deep dives, detailed FAQs) is published without dense Schema, without FAQPage markup, without proper Article markup. A 1,500-word article without Schema is worth less to an LLM than a 200-word Wikipedia entry. Textual density does not compensate for the absence of machine structure.

6. What an operator can activate in 90 days

The PROEMA operational sequence for this vertical fits into three 30-day sprints. Sprint 1 (days 1 to 30): audit the existing graph (Schema, robots.txt, llms.txt, Wikidata, sitemap), complete Wikidata with the vertical’s most relevant external identifiers, deploy dense Schema.org markup on the home and strategic pages. Sprint 2 (days 31 to 60): rebuild category pages (“where to buy”, “who to choose”), add a structured FAQ with FAQPage Schema, translate or natively write the NL and EN versions. Sprint 3 (days 61 to 90): monitor citation rate on a panel of 30 to 50 target queries, adjust editorially, add a tier-1 press block with sameAs Wikipedia.

The order-of-magnitude cost for this sequence is between EUR 12,000 and EUR 22,000, excluding internal validation time. The typical gain observed across the GEO Rocket portfolio on comparable verticals ranges from +11 to +21 points of Perplexity citation rate in 90 days, with persistence beyond 12 months. The first-mover premium then remains defendable for an additional 12 to 18 months if direct competitors do not structure within the same window.

The 35-point PROEMA GEO checklist was deliberately published open source inside PGSM v1.0. No comparable method existed before May 2026: MentionLab, Profound and Peec show aggregated scores with no sector reading grid and no public source for their method.

7. Common pitfalls to avoid

Three recurring pitfalls hit operators who launch a GEO strategy without structured guidance. First: confusing SEO and GEO. SEO optimizes for Google Search through keywords and backlinks. GEO optimizes for LLMs through machine structure and Wikidata authority signals. An SEO consultant who repackages an offering as “GEO” without changing methodology produces marginal results. PROEMA’s Q1 2026 study on 80 comparative audits shows a citation rate delta of only 4 to 7 points for “repackaged SEO” approaches, versus 11 to 21 points for GEO-native approaches.

Second: full content outsourcing to LLMs without human review. This practice, paradoxically widespread since 2024, is doubly counterproductive in 2026. On one hand, the EU AI Act Article 50 (effective 2 August 2026) requires explicit labeling of AI content, which makes non-compliance legally risky. On the other hand, mainstream LLMs are starting to negatively weight sources that exhibit strong markers of unreviewed automatic generation (stylistic homogeneity, lexical redundancy, no verifiable human signature). GEO 2026 rewards human or reviewed hybrid prose, not raw machine output.

Third: abandoning maintenance after sprint 3. A client who stops monitoring and Wikidata updates at day 90 loses on average 4 to 7 citation rate points over the following 12 months, through mechanical drift as competitors structure. GEO is not a one-off project. It is a continuous maintenance discipline, like treasury management or compliance.

8. What the next 12 months look like

The next twelve months will see three structural shifts in the GEO market. First: the consolidation of the AI Mode rollout in Europe (gradual deployment Q3 2026 through mid-2027), which will compress the position-1 SEO CTR from approximately 20% to 10-12% on informational queries. Operators who are not cited inside the AI Mode answer disappear from the value flow. Those who are cited capture a disproportionate share of residual value.

Second: the entry into force of the EU AI Act Article 50 on 2 August 2026, which makes machine-readable labeling of AI-generated content mandatory for any operator targeting EU users. This is not just an administrative cost. It is a labeling discipline that feeds into the machine confidence that LLMs grant to sources. A compliant brand gains methodological authority; a non-compliant one exposes itself to both fines (up to EUR 15M or 3% of global turnover) and a progressive machine confidence downgrade over the following months.

Third: the maturation of the third-party tracking tools market (MentionLab, Profound, Peec) into commoditization. Once these tools are commoditized, the real competitive differentiator becomes the agency methodology behind them: published, sourced, reproducible scoring like PGSM v1.0, or opaque agency-bound scoring. Brands that already work with method-driven agencies in 2026 enter 2027 with a defensible structural advantage.

9. How to decide

The decision to invest now hinges on three variables: annualized citation rate loss if nothing changes (4 to 6 points per year, mechanical, driven by structured competitors), opportunity cost on transactional queries (where to buy, price, distributor, hours) that capture margin directly, and the first-mover window on cultural and comparative queries still unsaturated in this vertical. On all three variables, the trade-off is clear: waiting is more expensive than acting.

The operator who structures their GEO in 2026 is not making a communications bet. They are protecting a brand asset built over years against an algorithmic drift that, without intervention, transfers narrative value to Wikipedia, to Anglo-Saxon aggregators, and to competitors who will have structured their graph first. The cost of waiting doubles every 9 to 12 months.

  • RFC 9309 (“Robots Exclusion Protocol”, IETF, September 2022)
  • llmstxt.org specification published by Jeremy Howard, September 2024
  • Princeton, arXiv 2311.09735, “GEO: Generative Engine Optimization”, November 2023
  • schema.org documentation (Article, FAQPage, BreadcrumbList, Organization, Person, Place)

The PROEMA GEO diagnostic is free: thirty minutes on one URL and three target prompts, answered within 24 hours. Gap reading in less than fifteen minutes. To go further, the PGSM v1.0 standard documents the full method as open source, and the method page details the five GEO levers applied to the GEO Rocket portfolio.

Get started

Is your brand citable by AI?

Receive your free GEO score and then request a full GEO diagnostic: what ChatGPT, Perplexity and Gemini do (or don’t) see about you today.