Betterfolio
Betterfolio team
GEO and AEO for IT services firms: appearing in ChatGPT, Perplexity and Copilot answers
Definitions, AI crawlers, page and competency dossier structure, a 30-day measurement plan. What an IT services firm can control, and what no provider can promise.
In short
An AI assistant only cites pages it can find, read and attribute. When a CIO asks ChatGPT or Perplexity which firm can staff a Java team in Lyon, the assistant runs searches, reads a handful of pages and writes an answer with its sources. A firm whose site blocks search crawlers, or whose pages contain no verifiable fact, does not appear in that answer.
- GEO (Generative Engine Optimization) is about a source being present in an answer written by a model. AEO (Answer Engine Optimization) is about the shape of the content: a passage an engine can reuse as-is as an answer.
- Both rest on SEO. Google states that no specific optimization is required for AI Overviews and AI Mode beyond standard Search best practices.
- First check: robots.txt and the firewall. OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot and Bingbot must reach the service pages. Blocking GPTBot, OpenAI's training crawler, does not remove a site from ChatGPT Search.
- Citable content carries dated, consistent facts: headcount, stacks, locations, staffing lead time, anonymized references. The same figure on every page.
- The 30-day measurement combines a fixed set of questions, the GA4 AI Assistant channel, the Bing Webmaster Tools AI Performance report and server logs.
- No citation is guaranteed. OpenAI states that placement in ChatGPT is not guaranteed. A provider promising a ranking inside an assistant is selling an outcome nobody controls.
GEO and AEO: two distinct ideas on the same foundation
Definitions
AEO, Answer Engine Optimization. The term spread with Google featured snippets and voice assistants, before generative models. The goal is for an engine to reuse a passage from the page as a direct answer to a question. The unit of work is the passage: a definition, a two-sentence answer, a table row.
GEO, Generative Engine Optimization. Pranjal Aggarwal and co-authors formalized the term in a paper published at KDD 2024. The goal is for a source to be reused, cited or mentioned in an answer a model writes from several pages. The unit of work is the whole source: whether it appears in the answer, where, and how much of the text draws on it.
For an IT services firm, the distinction is practical. AEO concerns how each page is written. GEO answers a sales leadership question: does our name appear in the answer, and is our offer described correctly?
SEO, AEO, GEO: what changes
| SEO | AEO | GEO | |
|---|---|---|---|
| Goal | A rank on a results page | A passage reused as the answer | A mention or citation in a generated answer |
| Surfaces | Google, Bing | Featured snippets, AI Overviews, voice assistants | ChatGPT, Perplexity, Copilot, Claude, Gemini, AI Mode |
| Unit of work | The page | The paragraph, the table row | The source and the entity, meaning the firm's name |
| Measurement question | Which position? | Which passage is reused? | Are we cited, and accurately? |
| Typical failure | Page 2 | Text too vague to extract | Absence, or an inaccurate description of the offer |
What the platforms say
Google states that SEO best practices still apply to AI Overviews and AI Mode, with no additional technical requirement. A page must be indexed and eligible for a snippet in Search to appear as a supporting link. Google adds that no AI-specific file and no special schema.org markup is needed.
OpenAI states that ChatGPT ranks search results using multiple factors and that placement is not guaranteed. To be eligible, OpenAI asks site owners to allow OAI-SearchBot and to confirm that the host or CDN accepts its published IP addresses.
GEO and AEO therefore do not replace SEO. They add two requirements to existing work: content written to be extracted, and a way to measure presence in answers.
Why IT services firms are affected
Two uses of assistants bear directly on selling services.
Pre-selection before the first contact
A CIO, a buyer or a project lead types the need into an assistant: "IT services firm able to staff four Java Spring developers in Lyon, start within a month." OpenAI documents what happens next in its search help article. ChatGPT rewrites the request into one or more short queries, sends them to partner search providers, reviews the results and writes its answer. It also passes along an approximate location for the user.
Two consequences follow. The query actually sent looks like keywords, such as "IT services Java Spring Lyon", so service pages must still rank well. The final answer reuses precise sentences, so those pages must contain facts the assistant can quote without distorting them.
The client's assistant reading your dossiers
A buyer who receives three competency dossiers can drop them into an assistant and ask for a comparison. The dossier is then read by a system before a person reads it. A scanned PDF with no headings and 25 technologies in a pile produces a vague summary. A structured dossier produces an accurate one.
This second use has nothing to do with search rankings. It is a matter of document quality, and it can be fixed without waiting for any engine to index anything.
How an assistant selects its sources
For answers with links, assistants do not rely primarily on training data. They query a search index at the time of the question, their own or a partner's, then read the pages they select. A page missing from those indexes cannot be cited.
Three families of crawlers
The main vendors separate three uses, with a named crawler for each.
| Family | Role | Effect of blocking it |
|---|---|---|
| Search crawler | Builds the index used to answer | The site drops out of the assistant's search answers |
| User-action fetcher | Reads a page in real time at a user's request | The assistant can no longer open the page when asked to |
| Training crawler | Collects content to train models | No effect on search. The content leaves future training sets |
The crawlers to know
| Vendor | Search, to allow | User action | Training, a policy choice |
|---|---|---|---|
| OpenAI (ChatGPT) | OAI-SearchBot | ChatGPT-User | GPTBot |
| Anthropic (Claude) | Claude-SearchBot | Claude-User | ClaudeBot |
| Perplexity | PerplexityBot | Perplexity-User | No declared crawler |
| Mistral | MistralAI-Index | MistralAI-User | MistralAI-Training |
| Google (AI Overviews, AI Mode, Gemini) | Googlebot | No dedicated fetcher | Google-Extended, a control token |
| Microsoft (Copilot) | Bingbot | No dedicated fetcher | Not documented |
Three points from the official documentation:
- ChatGPT-User and Perplexity-User act on a user's behalf. OpenAI says robots.txt rules may not apply to ChatGPT-User. Perplexity says Perplexity-User generally ignores them. Presence in search is controlled through OAI-SearchBot and PerplexityBot.
- Google-Extended is not a crawler but a control token. It has no effect on Google Search or AI Overviews. It governs whether content is used to train Gemini and to ground answers in the Gemini app. Blocking it can reduce presence in Gemini, not in Google.
- Blocking GPTBot or ClaudeBot is a training policy decision. It removes the site from neither ChatGPT Search nor Claude. The mistake that costs visibility is the opposite one: a blanket "block AI" rule that also shuts out search crawlers.
Step 1: check technical access
Before rewriting anything, run four checks. They take the person who manages the site, not a redesign.
The robots.txt file
Open https://your-domain.com/robots.txt. Two situations cause problems: a Disallow: / under User-agent: * with no exception for search crawlers, or a Disallow under the name of a search crawler.
Example configuration for a firm that accepts search and declines training:
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: MistralAI-Index
Allow: /
Disallow: /client-area/
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: MistralAI-Training
Disallow: /
User-agent: *
Allow: /
Disallow: /client-area/
One technical rule prevents many errors. Under RFC 9309, a crawler that finds a group addressed to its name applies that group and ignores User-agent: *. General exclusions, such as a client area, must therefore be repeated in every named group. OpenAI and Perplexity say a change to the file can take up to 24 hours to take effect.
Declining training remains a choice. It has no effect on visibility in answers, in either direction.
The CDN and the firewall
An open robots.txt is not enough if the firewall rejects the requests. OpenAI asks site owners to confirm that the host or CDN accepts the IP addresses published for OAI-SearchBot. Perplexity documents the same steps for Cloudflare and AWS WAF.
In July 2026 Cloudflare announced a classification of AI traffic into three separately configurable categories: Search, Agent and Training. The new defaults have applied since September 15, 2026. According to Cloudflare's documentation, a Training block also affects mixed-purpose crawlers that serve both search and training, including through the legacy "Block AI bots" option. A firm that turned this option on should check what it blocks today, under Security Settings, Configure AI bot policies.
If Cloudflare's managed robots.txt is enabled, Cloudflare adds its own directives to the file. Check the file actually served online, not the one in the code repository.
HTML rendering
In December 2024 Vercel and the consultancy MERJ published an analysis of AI crawler traffic on Vercel's network. At that date, none of the major AI crawlers executed JavaScript: OAI-SearchBot, ChatGPT-User, GPTBot, ClaudeBot and PerplexityBot read the initial HTML. Googlebot, which also serves Gemini, renders pages.
A site whose service descriptions only appear in the browser therefore looks empty to those crawlers. The test takes a minute. View the page source (Ctrl+U) and search for a sentence from the offer. If it is not there, AI crawlers probably do not read it.
Google and Bing indexing
Google's AI answers start from the Google index. Copilot relies on the Bing index. ChatGPT can obtain URLs through partner search providers as well as its own crawler. Check in Google Search Console and Bing Webmaster Tools that the service pages are indexed.
Registering with Bing Webmaster Tools is worth it even when Bing brings little traffic. Its AI Performance report is, to our knowledge, the only vendor-provided tool that counts a site's citations in AI answers.
Summary of checks
| Check | Where | Expected result |
|---|---|---|
| Served robots.txt | your-domain.com/robots.txt | No Disallow on search crawlers |
| Firewall, CDN | Cloudflare or hosting console | Search category allowed, no Training block that catches mixed-purpose crawlers |
| Text in the HTML | Page source | The offer's facts appear without JavaScript |
| Google indexing | Search Console, URL inspection | Service pages indexed, snippets allowed |
| Bing indexing | Bing Webmaster Tools | Service pages indexed, sitemap submitted |
Step 2: write answer-ready pages
An answer-ready page answers a buying question in its first lines, with facts an assistant can reuse on their own. A persuasive tone does not help. In the GEO study, a more assertive style brought no significant visibility gain. Adding sources, figures and quotations did.
Formats that extract well
| Format | Use | Example for an IT services firm |
|---|---|---|
| In short | Three to five lines at the top of the page that answer on their own | "35 Java and Spring consultants, Lyon and Grenoble, staffing within 3 weeks" |
| Definition | One sentence that fixes a term | "Mid-level: 4 to 7 years, autonomous on a scope, no cross-cutting architecture" |
| Table | Compare, list criteria, set out a grid | Stacks, number of consultants, sectors already staffed |
| FAQ | Reuse buyers' questions in their own words | "How quickly can you staff a team of four?" |
| Dated proof | An anonymized, quantified, dated engagement | "2025, insurance, Spring monolith migrated to EKS, team of 4" |
Five writing rules
- The first sentence states the fact. Not the intention, not the vision. "We staff Java teams of 2 to 8 people within 3 weeks in the Auvergne-Rhône-Alpes region."
- Every paragraph stays accurate when quoted alone. No "as mentioned above", no pronoun pointing back to the previous section.
- Subheadings reuse buying questions. "Staffing lead time" beats "Our responsiveness".
- Figures are dated and conservative. Actual 2026 headcount rather than a cumulative total since founding.
- Every strong claim has its source. A certification with its issuing body, a reference with its year, a study with its author.
The same page, two versions
Non-extractable version: "We support CIOs in their cloud transformation with a team of passionate experts who integrate quickly into your teams."
Answer-ready version: "Lyon-based IT services firm with 120 consultants. 35 Java 17 and Spring Boot profiles available within 3 weeks, 11 of them with an engagement of more than six months in banking or insurance. Latest team delivered: 4 people, migration of a Spring monolith to services on Amazon EKS, insurance sector, 2025."
The second version contains no adjectives. It contains six verifiable facts: city, headcount, stack, lead time, sector, dated reference. An assistant can reuse any one of them without distorting the offer.
Entity consistency
A model that assembles several sources needs to identify the company without ambiguity.
- One name. Not "Dupont IT" on the website, "Dupont Consulting" on LinkedIn and "Dupont Group" in tender responses. To a reader, it is the same company. To a model, it is three entities.
- The same figures everywhere. If the Cloud page claims 80 AWS experts and the About page 50 cloud consultants, the assistant holds two contradictory facts. It keeps one, or neither.
- The same labels. Same stack names, same seniority levels, on the website, in dossiers and on consultants' public profiles.
What the research shows, and its limits
The GEO paper by Aggarwal et al. tested nine rewriting methods on GEO-bench, a benchmark of 10,000 queries across 25 domains. The three most effective methods added cited sources, quotations and statistics. They raised visibility by 30 to 40% on the study's main metric. Keyword stuffing brought little or no gain. On Perplexity, the authors measured gains of up to 37%.
Three caveats for an IT services firm:
- These gains are measured maximums under a specific protocol. They are not an expected average.
- They concentrate on sources that were poorly ranked to begin with. For sources already in first position, adding cited sources lowered visibility by about 30% in the published results.
- The study covers general-purpose queries. It does not measure B2B staffing queries.
The usable lesson is modest. Sourced, quantified facts help. A sales tone does not.
Structured data: useful, not decisive
Schema.org markup, such as Organization, Article or FAQPage, helps engines interpret a page. It does not trigger citations.
- Since August 2023, Google only shows FAQ rich results for well-known government and health websites. FAQ markup on an IT services site will not produce a rich result.
- Google states that no special markup is required for its AI features, and that structured data must match the visible text.
The FAQ keeps its value for another reason. It puts the question and its answer side by side, which makes it the easiest format to extract, for an assistant and for a reader.
Step 3: treat the competency dossier as a source
A dossier sent by email sits outside any search index. It still becomes a source the moment the buyer drops it into an assistant. Its structure determines how faithful the summary is.
Building blocks of an extractable dossier
| Block | Content | What the assistant takes from it |
|---|---|---|
| Header | Role, seniority, city, availability | The profile's identity and the staffing constraint |
| Priority stack | 6 to 8 technologies, level, last year of use | The match with the need |
| Three key engagements | Context, actual role, duration, deliverable | The proof, with its context |
| Led or contributed | What the consultant decided, what they worked alongside | A boundary that prevents inflated attributions |
| Sectors | Banking, insurance, public sector | A credibility filter |
| Out of scope | What the profile does not cover | Less room for extrapolation |
The distinction between led and contributed has a direct effect on automated summaries. A mid-level developer who worked in a team using Kubernetes should not become a "Kubernetes expert" in the summary produced on the client side. Stated explicitly, the boundary leaves less room for invention.
Three checks before sending
- The PDF contains selectable text. A scan or an image of text forces the assistant through character recognition, with errors.
- Headings are real headings. A document with heading levels splits correctly. Text bolded by hand splits less well.
- The vocabulary matches the website. Same job title, same stack names, same seniority scale.
One test per template is enough. Drop an anonymized dossier into ChatGPT or Claude, ask "Summarize this profile in five lines and list what it does not cover", then compare with what the firm would have written. Each gap points to an ambiguity in the document. Anonymization is not a matter of style: a named dossier contains personal data, and sending it to a third-party service falls under the GDPR.
Measuring over 30 days
Assistant answers vary from one attempt to the next, depending on wording, account and location. What you measure is a frequency over a fixed set of questions, not a position.
Available measurement sources
| Source | What it measures | Limit |
|---|---|---|
| Question panel, in a spreadsheet | Mention, citation with link, accuracy, per assistant | Manual, small sample, variable answers |
| GA4, AI Assistant default channel | Visits from ChatGPT, Gemini, Copilot, DeepSeek and Grok | Perplexity not on Google's list, not retroactive, visits without a referrer not counted |
| GA4, custom channel group | Every AI source defined by a rule, Perplexity included | Must sit above "Referral" in the channel order |
utm_source=chatgpt.com parameter | Clicks from ChatGPT search results | Missing when the link is copied and pasted elsewhere |
| Bing Webmaster Tools, AI Performance | Citations in Copilot and Bing AI summaries, by URL, with grounding queries | Microsoft ecosystem only, preview feature |
| Server logs | Visits from OAI-SearchBot, PerplexityBot, Claude-SearchBot, ChatGPT-User | Requires access to the host's logs |
| Google Search Console | Search traffic, AI Overviews and AI Mode included | No filter separates AI answers from the rest |
A visit from ChatGPT-User or Perplexity-User in the logs means a user asked a question that led the assistant to open the page. It is the most direct signal of real consultation.
The week-by-week plan
Week 1: access and reference facts
- Check robots.txt, firewall, HTML rendering, Google and Bing indexing
- Settle the company name and a fact sheet: headcount, main stacks, locations, staffing lead time, sectors
- Create the AI channel group in GA4 and register the site in Bing Webmaster Tools
- Write the panel: 15 questions clients actually ask, phrased the way they phrase them
Week 2: baseline
- Ask the 15 questions in ChatGPT, Perplexity, Copilot and Gemini on the same day, with no conversation history
- Log for each answer the firms cited, whether a link is present and whether the description is accurate
- Pull AI crawler visits from the last 30 days of server logs
Week 3: fixes
- Rewrite the three most requested service pages with an In short block, an FAQ and two dated proofs
- Align the About page and service pages on the same figures
- Move competency dossiers to the structured template and test an anonymized dossier in an assistant
Week 4: second reading
- Ask the same 15 questions again, without rewording them
- Compare appearances, disappearances, and descriptions that were corrected or distorted
- Deal with inaccurate answers first. A wrong description of the offer weighs more than an absence
The tracking log
| Question | Assistant | Date | Firms cited | Our firm cited | With link | Accurate description |
|---|---|---|---|---|---|---|
| Java Spring firm Lyon, 4 people, within a month | Perplexity | 2026-10-06 | A, C, D | No | N/A | N/A |
| Java Spring firm Lyon, 4 people, within a month | ChatGPT | 2026-10-06 | B, D | Yes (B) | Yes | No, 2023 headcount |
Three indicators are enough: mention rate, citation-with-link rate and accuracy rate. With 15 questions and four assistants, that is 60 readings, and a difference of two or three answers between two months is not enough to draw a conclusion.
What 30 days can tell you
One month gives a first inventory, not a result. OpenAI and Perplexity pick up a modified robots.txt in about 24 hours. At Google, recrawling a modified page can take anywhere from several days to several months. The reading becomes meaningful when it is repeated monthly over a quarter.
Limits to keep in mind
- Citation is not guaranteed. OpenAI says so for ChatGPT. Google says so for indexing and serving. No provider controls an assistant's final choice.
- Answers vary. Wording, history, location and model version change the list. A single reading proves nothing.
- An llms.txt file is not a Google lever. Google states that no AI-specific file is needed for its AI features. Other tools may read it. It replaces neither indexing nor published facts.
- Measurement undercounts. Some visits from assistants arrive without a referrer and land in Direct in GA4. The figures you get are a floor.
- An inflated figure backfires. The assistant repeats the published figure, and the buyer checks it in the interview.
- GEO does not replace the sales relationship. A citation opens a first contact. The dossier, the conversation and the staffing remain the job.
Frequently asked questions
What is the difference between GEO and AEO?
AEO (Answer Engine Optimization) works on the shape of content so an engine reuses a passage as a direct answer: definition, short answer, table, FAQ. GEO (Generative Engine Optimization) targets the presence of a source in an answer a model writes from several pages, and the accuracy of what it says about that source. Both rely on sound SEO.
Should we block GPTBot to protect our content?
It is a training policy choice with no effect on visibility in ChatGPT Search. OpenAI documents two independent crawlers: GPTBot for model training and OAI-SearchBot for search. A firm can decline training and remain visible in answers, provided it does not block OAI-SearchBot. The same separation exists at Anthropic (ClaudeBot and Claude-SearchBot) and at Mistral (MistralAI-Training and MistralAI-Index).
Can a citation in ChatGPT or Perplexity be guaranteed?
No. OpenAI states that ChatGPT ranks results using multiple factors and that placement is not guaranteed. Allowing search crawlers makes a page eligible, nothing more. The right measure is a mention rate over a fixed set of questions, taken every month.
Is an llms.txt file necessary?
Not for Google, which states that no AI-specific file is required for AI Overviews and AI Mode. Other tools may read the file. It replaces neither search crawler access nor factual service pages, which remain the priority.
How can we tell whether ChatGPT sends visits to our site?
In GA4, the AI Assistant default channel has grouped visits from ChatGPT, Gemini, Copilot, DeepSeek and Grok since May 2026. ChatGPT also adds the utm_source=chatgpt.com parameter to links in its search results. For Perplexity, you need a custom channel group. These figures remain a floor, because some visits arrive without a referrer.
Where should we start with two weeks available?
Check robots.txt and the firewall, settle a single fact sheet (name, headcount, stacks, locations, lead time), rewrite the most requested service page with an In short block and an FAQ, and move ten dossiers to the structured template. Then ask ten client questions in ChatGPT and Perplexity and log the answers. The rest of the plan can follow the next month.
Structuring dossiers without turning it into an SEO project
A dossier generator does not place a firm inside ChatGPT. It addresses a frequent cause of poorly summarized dossiers: lack of time, which leads teams to resend an old Word document updated in a hurry.
| Step | What the tool handles | What the firm decides |
|---|---|---|
| Template | The same sections from one dossier to the next: role, skills, engagements, sectors | Which three engagements to feature |
| Engagement write-up | Consistent layout for context, role and deliverable | Where to draw the line between led and contributed |
| Vocabulary | The same labels and levels across every dossier | The volumes and lead times actually delivered |
| Export | Generated PDF with selectable text and a heading hierarchy | Sending it to the client |
The dossier that reaches the buyer is the one their assistant will summarize. If it is structured, the summary is more likely to be accurate.
Key takeaways
GEO is about being cited, and accurately, in a generated answer. AEO is about writing passages an engine can reuse on their own. Neither replaces SEO, and nobody can guarantee a citation.
Three workstreams for an IT services firm:
- Access. robots.txt, firewall and HTML rendering let OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot and Bingbot through. Declining training remains a separate decision.
- Facts. Every service page opens with a quantified In short block, answers buying questions in an FAQ and cites dated proof. The same figures appear on the website, in dossiers and in tender responses.
- Measurement. Fifteen fixed questions, four assistants, one reading a month, supplemented by GA4, Bing Webmaster Tools and server logs. Inaccurate answers get fixed before absences.
An assistant repeats what the firm publishes. The work consists of publishing accurate facts that a machine and a buyer can both read.
Sources
- Google Search Central, AI features and your website
- Google Search Central, List of Google's common crawlers, Google-Extended section
- Google Search Central, Changes to HowTo and FAQ rich results, August 2023
- OpenAI, Overview of OpenAI Crawlers
- OpenAI Help Center, Searching the web with ChatGPT
- OpenAI Help Center, Publishers and Developers FAQ
- Anthropic, Does Anthropic crawl data from the web?
- Perplexity, Perplexity Crawlers
- Mistral AI, Mistral crawlers
- Cloudflare, New options to manage AI traffic, July 2026
- Google Analytics, Default channel group
- Microsoft Bing, Introducing AI Performance in Bing Webmaster Tools, February 2026
- Vercel and MERJ, The rise of the AI crawler, December 2024
- Aggarwal et al., GEO: Generative Engine Optimization, KDD 2024
- IETF, RFC 9309, Robots Exclusion Protocol
Read next
Betterfolio vs Showcase (DoYouBuzz) vs Middleman: which tool for your IT services firm in 2026?
A comparison of the three French competency portfolio tools. BoondManager integration, real day-to-day adoption by managers, candidate qualification and pricing models: what actually matters when you choose.
The first 30 days on mission: securing the placement after the signature
A signed contract isn't the end of the sales cycle. Most missions that end early are decided in the first month. An onboarding protocol that protects margin and the client relationship.
Enterprise RFPs: industrialising your response without standardising the profiles
Answering an RFP fast and answering it well are contradictory demands, unless what you standardise is the process, not the content. A method for holding both.
