Betterfolio
Betterfolio

Betterfolio

Betterfolio team

GEO and AEO for IT services firms: appearing in ChatGPT, Perplexity and Copilot answers

Definitions, AI crawlers, page and competency dossier structure, a 30-day measurement plan. What an IT services firm can control, and what no provider can promise.

Laptop open on a code editor screen, Unsplash photo

In short

An AI assistant only cites pages it can find, read and attribute. When a CIO asks ChatGPT or Perplexity which firm can staff a Java team in Lyon, the assistant runs searches, reads a handful of pages and writes an answer with its sources. A firm whose site blocks search crawlers, or whose pages contain no verifiable fact, does not appear in that answer.

  • GEO (Generative Engine Optimization) is about a source being present in an answer written by a model. AEO (Answer Engine Optimization) is about the shape of the content: a passage an engine can reuse as-is as an answer.
  • Both rest on SEO. Google states that no specific optimization is required for AI Overviews and AI Mode beyond standard Search best practices.
  • First check: robots.txt and the firewall. OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot and Bingbot must reach the service pages. Blocking GPTBot, OpenAI's training crawler, does not remove a site from ChatGPT Search.
  • Citable content carries dated, consistent facts: headcount, stacks, locations, staffing lead time, anonymized references. The same figure on every page.
  • The 30-day measurement combines a fixed set of questions, the GA4 AI Assistant channel, the Bing Webmaster Tools AI Performance report and server logs.
  • No citation is guaranteed. OpenAI states that placement in ChatGPT is not guaranteed. A provider promising a ranking inside an assistant is selling an outcome nobody controls.

GEO and AEO: two distinct ideas on the same foundation

Definitions

AEO, Answer Engine Optimization. The term spread with Google featured snippets and voice assistants, before generative models. The goal is for an engine to reuse a passage from the page as a direct answer to a question. The unit of work is the passage: a definition, a two-sentence answer, a table row.

GEO, Generative Engine Optimization. Pranjal Aggarwal and co-authors formalized the term in a paper published at KDD 2024. The goal is for a source to be reused, cited or mentioned in an answer a model writes from several pages. The unit of work is the whole source: whether it appears in the answer, where, and how much of the text draws on it.

For an IT services firm, the distinction is practical. AEO concerns how each page is written. GEO answers a sales leadership question: does our name appear in the answer, and is our offer described correctly?

SEO, AEO, GEO: what changes

SEOAEOGEO
GoalA rank on a results pageA passage reused as the answerA mention or citation in a generated answer
SurfacesGoogle, BingFeatured snippets, AI Overviews, voice assistantsChatGPT, Perplexity, Copilot, Claude, Gemini, AI Mode
Unit of workThe pageThe paragraph, the table rowThe source and the entity, meaning the firm's name
Measurement questionWhich position?Which passage is reused?Are we cited, and accurately?
Typical failurePage 2Text too vague to extractAbsence, or an inaccurate description of the offer

What the platforms say

Google states that SEO best practices still apply to AI Overviews and AI Mode, with no additional technical requirement. A page must be indexed and eligible for a snippet in Search to appear as a supporting link. Google adds that no AI-specific file and no special schema.org markup is needed.

OpenAI states that ChatGPT ranks search results using multiple factors and that placement is not guaranteed. To be eligible, OpenAI asks site owners to allow OAI-SearchBot and to confirm that the host or CDN accepts its published IP addresses.

GEO and AEO therefore do not replace SEO. They add two requirements to existing work: content written to be extracted, and a way to measure presence in answers.


Why IT services firms are affected

Two uses of assistants bear directly on selling services.

Pre-selection before the first contact

A CIO, a buyer or a project lead types the need into an assistant: "IT services firm able to staff four Java Spring developers in Lyon, start within a month." OpenAI documents what happens next in its search help article. ChatGPT rewrites the request into one or more short queries, sends them to partner search providers, reviews the results and writes its answer. It also passes along an approximate location for the user.

Two consequences follow. The query actually sent looks like keywords, such as "IT services Java Spring Lyon", so service pages must still rank well. The final answer reuses precise sentences, so those pages must contain facts the assistant can quote without distorting them.

The client's assistant reading your dossiers

A buyer who receives three competency dossiers can drop them into an assistant and ask for a comparison. The dossier is then read by a system before a person reads it. A scanned PDF with no headings and 25 technologies in a pile produces a vague summary. A structured dossier produces an accurate one.

This second use has nothing to do with search rankings. It is a matter of document quality, and it can be fixed without waiting for any engine to index anything.


How an assistant selects its sources

For answers with links, assistants do not rely primarily on training data. They query a search index at the time of the question, their own or a partner's, then read the pages they select. A page missing from those indexes cannot be cited.

Three families of crawlers

The main vendors separate three uses, with a named crawler for each.

FamilyRoleEffect of blocking it
Search crawlerBuilds the index used to answerThe site drops out of the assistant's search answers
User-action fetcherReads a page in real time at a user's requestThe assistant can no longer open the page when asked to
Training crawlerCollects content to train modelsNo effect on search. The content leaves future training sets

The crawlers to know

VendorSearch, to allowUser actionTraining, a policy choice
OpenAI (ChatGPT)OAI-SearchBotChatGPT-UserGPTBot
Anthropic (Claude)Claude-SearchBotClaude-UserClaudeBot
PerplexityPerplexityBotPerplexity-UserNo declared crawler
MistralMistralAI-IndexMistralAI-UserMistralAI-Training
Google (AI Overviews, AI Mode, Gemini)GooglebotNo dedicated fetcherGoogle-Extended, a control token
Microsoft (Copilot)BingbotNo dedicated fetcherNot documented

Three points from the official documentation:

  • ChatGPT-User and Perplexity-User act on a user's behalf. OpenAI says robots.txt rules may not apply to ChatGPT-User. Perplexity says Perplexity-User generally ignores them. Presence in search is controlled through OAI-SearchBot and PerplexityBot.
  • Google-Extended is not a crawler but a control token. It has no effect on Google Search or AI Overviews. It governs whether content is used to train Gemini and to ground answers in the Gemini app. Blocking it can reduce presence in Gemini, not in Google.
  • Blocking GPTBot or ClaudeBot is a training policy decision. It removes the site from neither ChatGPT Search nor Claude. The mistake that costs visibility is the opposite one: a blanket "block AI" rule that also shuts out search crawlers.

Step 1: check technical access

Before rewriting anything, run four checks. They take the person who manages the site, not a redesign.

The robots.txt file

Open https://your-domain.com/robots.txt. Two situations cause problems: a Disallow: / under User-agent: * with no exception for search crawlers, or a Disallow under the name of a search crawler.

Example configuration for a firm that accepts search and declines training:

User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: MistralAI-Index
Allow: /
Disallow: /client-area/

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: MistralAI-Training
Disallow: /

User-agent: *
Allow: /
Disallow: /client-area/

One technical rule prevents many errors. Under RFC 9309, a crawler that finds a group addressed to its name applies that group and ignores User-agent: *. General exclusions, such as a client area, must therefore be repeated in every named group. OpenAI and Perplexity say a change to the file can take up to 24 hours to take effect.

Declining training remains a choice. It has no effect on visibility in answers, in either direction.

The CDN and the firewall

An open robots.txt is not enough if the firewall rejects the requests. OpenAI asks site owners to confirm that the host or CDN accepts the IP addresses published for OAI-SearchBot. Perplexity documents the same steps for Cloudflare and AWS WAF.

In July 2026 Cloudflare announced a classification of AI traffic into three separately configurable categories: Search, Agent and Training. The new defaults have applied since September 15, 2026. According to Cloudflare's documentation, a Training block also affects mixed-purpose crawlers that serve both search and training, including through the legacy "Block AI bots" option. A firm that turned this option on should check what it blocks today, under Security Settings, Configure AI bot policies.

If Cloudflare's managed robots.txt is enabled, Cloudflare adds its own directives to the file. Check the file actually served online, not the one in the code repository.

HTML rendering

In December 2024 Vercel and the consultancy MERJ published an analysis of AI crawler traffic on Vercel's network. At that date, none of the major AI crawlers executed JavaScript: OAI-SearchBot, ChatGPT-User, GPTBot, ClaudeBot and PerplexityBot read the initial HTML. Googlebot, which also serves Gemini, renders pages.

A site whose service descriptions only appear in the browser therefore looks empty to those crawlers. The test takes a minute. View the page source (Ctrl+U) and search for a sentence from the offer. If it is not there, AI crawlers probably do not read it.

Google and Bing indexing

Google's AI answers start from the Google index. Copilot relies on the Bing index. ChatGPT can obtain URLs through partner search providers as well as its own crawler. Check in Google Search Console and Bing Webmaster Tools that the service pages are indexed.

Registering with Bing Webmaster Tools is worth it even when Bing brings little traffic. Its AI Performance report is, to our knowledge, the only vendor-provided tool that counts a site's citations in AI answers.

Summary of checks

CheckWhereExpected result
Served robots.txtyour-domain.com/robots.txtNo Disallow on search crawlers
Firewall, CDNCloudflare or hosting consoleSearch category allowed, no Training block that catches mixed-purpose crawlers
Text in the HTMLPage sourceThe offer's facts appear without JavaScript
Google indexingSearch Console, URL inspectionService pages indexed, snippets allowed
Bing indexingBing Webmaster ToolsService pages indexed, sitemap submitted

Step 2: write answer-ready pages

An answer-ready page answers a buying question in its first lines, with facts an assistant can reuse on their own. A persuasive tone does not help. In the GEO study, a more assertive style brought no significant visibility gain. Adding sources, figures and quotations did.

Formats that extract well

FormatUseExample for an IT services firm
In shortThree to five lines at the top of the page that answer on their own"35 Java and Spring consultants, Lyon and Grenoble, staffing within 3 weeks"
DefinitionOne sentence that fixes a term"Mid-level: 4 to 7 years, autonomous on a scope, no cross-cutting architecture"
TableCompare, list criteria, set out a gridStacks, number of consultants, sectors already staffed
FAQReuse buyers' questions in their own words"How quickly can you staff a team of four?"
Dated proofAn anonymized, quantified, dated engagement"2025, insurance, Spring monolith migrated to EKS, team of 4"

Five writing rules

  1. The first sentence states the fact. Not the intention, not the vision. "We staff Java teams of 2 to 8 people within 3 weeks in the Auvergne-Rhône-Alpes region."
  2. Every paragraph stays accurate when quoted alone. No "as mentioned above", no pronoun pointing back to the previous section.
  3. Subheadings reuse buying questions. "Staffing lead time" beats "Our responsiveness".
  4. Figures are dated and conservative. Actual 2026 headcount rather than a cumulative total since founding.
  5. Every strong claim has its source. A certification with its issuing body, a reference with its year, a study with its author.

The same page, two versions

Non-extractable version: "We support CIOs in their cloud transformation with a team of passionate experts who integrate quickly into your teams."

Answer-ready version: "Lyon-based IT services firm with 120 consultants. 35 Java 17 and Spring Boot profiles available within 3 weeks, 11 of them with an engagement of more than six months in banking or insurance. Latest team delivered: 4 people, migration of a Spring monolith to services on Amazon EKS, insurance sector, 2025."

The second version contains no adjectives. It contains six verifiable facts: city, headcount, stack, lead time, sector, dated reference. An assistant can reuse any one of them without distorting the offer.

Entity consistency

A model that assembles several sources needs to identify the company without ambiguity.

  • One name. Not "Dupont IT" on the website, "Dupont Consulting" on LinkedIn and "Dupont Group" in tender responses. To a reader, it is the same company. To a model, it is three entities.
  • The same figures everywhere. If the Cloud page claims 80 AWS experts and the About page 50 cloud consultants, the assistant holds two contradictory facts. It keeps one, or neither.
  • The same labels. Same stack names, same seniority levels, on the website, in dossiers and on consultants' public profiles.

What the research shows, and its limits

The GEO paper by Aggarwal et al. tested nine rewriting methods on GEO-bench, a benchmark of 10,000 queries across 25 domains. The three most effective methods added cited sources, quotations and statistics. They raised visibility by 30 to 40% on the study's main metric. Keyword stuffing brought little or no gain. On Perplexity, the authors measured gains of up to 37%.

Three caveats for an IT services firm:

  • These gains are measured maximums under a specific protocol. They are not an expected average.
  • They concentrate on sources that were poorly ranked to begin with. For sources already in first position, adding cited sources lowered visibility by about 30% in the published results.
  • The study covers general-purpose queries. It does not measure B2B staffing queries.

The usable lesson is modest. Sourced, quantified facts help. A sales tone does not.

Structured data: useful, not decisive

Schema.org markup, such as Organization, Article or FAQPage, helps engines interpret a page. It does not trigger citations.

  • Since August 2023, Google only shows FAQ rich results for well-known government and health websites. FAQ markup on an IT services site will not produce a rich result.
  • Google states that no special markup is required for its AI features, and that structured data must match the visible text.

The FAQ keeps its value for another reason. It puts the question and its answer side by side, which makes it the easiest format to extract, for an assistant and for a reader.


Step 3: treat the competency dossier as a source

A dossier sent by email sits outside any search index. It still becomes a source the moment the buyer drops it into an assistant. Its structure determines how faithful the summary is.

Building blocks of an extractable dossier

BlockContentWhat the assistant takes from it
HeaderRole, seniority, city, availabilityThe profile's identity and the staffing constraint
Priority stack6 to 8 technologies, level, last year of useThe match with the need
Three key engagementsContext, actual role, duration, deliverableThe proof, with its context
Led or contributedWhat the consultant decided, what they worked alongsideA boundary that prevents inflated attributions
SectorsBanking, insurance, public sectorA credibility filter
Out of scopeWhat the profile does not coverLess room for extrapolation

The distinction between led and contributed has a direct effect on automated summaries. A mid-level developer who worked in a team using Kubernetes should not become a "Kubernetes expert" in the summary produced on the client side. Stated explicitly, the boundary leaves less room for invention.

Three checks before sending

  1. The PDF contains selectable text. A scan or an image of text forces the assistant through character recognition, with errors.
  2. Headings are real headings. A document with heading levels splits correctly. Text bolded by hand splits less well.
  3. The vocabulary matches the website. Same job title, same stack names, same seniority scale.

One test per template is enough. Drop an anonymized dossier into ChatGPT or Claude, ask "Summarize this profile in five lines and list what it does not cover", then compare with what the firm would have written. Each gap points to an ambiguity in the document. Anonymization is not a matter of style: a named dossier contains personal data, and sending it to a third-party service falls under the GDPR.


Measuring over 30 days

Assistant answers vary from one attempt to the next, depending on wording, account and location. What you measure is a frequency over a fixed set of questions, not a position.

Available measurement sources

SourceWhat it measuresLimit
Question panel, in a spreadsheetMention, citation with link, accuracy, per assistantManual, small sample, variable answers
GA4, AI Assistant default channelVisits from ChatGPT, Gemini, Copilot, DeepSeek and GrokPerplexity not on Google's list, not retroactive, visits without a referrer not counted
GA4, custom channel groupEvery AI source defined by a rule, Perplexity includedMust sit above "Referral" in the channel order
utm_source=chatgpt.com parameterClicks from ChatGPT search resultsMissing when the link is copied and pasted elsewhere
Bing Webmaster Tools, AI PerformanceCitations in Copilot and Bing AI summaries, by URL, with grounding queriesMicrosoft ecosystem only, preview feature
Server logsVisits from OAI-SearchBot, PerplexityBot, Claude-SearchBot, ChatGPT-UserRequires access to the host's logs
Google Search ConsoleSearch traffic, AI Overviews and AI Mode includedNo filter separates AI answers from the rest

A visit from ChatGPT-User or Perplexity-User in the logs means a user asked a question that led the assistant to open the page. It is the most direct signal of real consultation.

The week-by-week plan

Week 1: access and reference facts

  1. Check robots.txt, firewall, HTML rendering, Google and Bing indexing
  2. Settle the company name and a fact sheet: headcount, main stacks, locations, staffing lead time, sectors
  3. Create the AI channel group in GA4 and register the site in Bing Webmaster Tools
  4. Write the panel: 15 questions clients actually ask, phrased the way they phrase them

Week 2: baseline

  1. Ask the 15 questions in ChatGPT, Perplexity, Copilot and Gemini on the same day, with no conversation history
  2. Log for each answer the firms cited, whether a link is present and whether the description is accurate
  3. Pull AI crawler visits from the last 30 days of server logs

Week 3: fixes

  1. Rewrite the three most requested service pages with an In short block, an FAQ and two dated proofs
  2. Align the About page and service pages on the same figures
  3. Move competency dossiers to the structured template and test an anonymized dossier in an assistant

Week 4: second reading

  1. Ask the same 15 questions again, without rewording them
  2. Compare appearances, disappearances, and descriptions that were corrected or distorted
  3. Deal with inaccurate answers first. A wrong description of the offer weighs more than an absence

The tracking log

QuestionAssistantDateFirms citedOur firm citedWith linkAccurate description
Java Spring firm Lyon, 4 people, within a monthPerplexity2026-10-06A, C, DNoN/AN/A
Java Spring firm Lyon, 4 people, within a monthChatGPT2026-10-06B, DYes (B)YesNo, 2023 headcount

Three indicators are enough: mention rate, citation-with-link rate and accuracy rate. With 15 questions and four assistants, that is 60 readings, and a difference of two or three answers between two months is not enough to draw a conclusion.

What 30 days can tell you

One month gives a first inventory, not a result. OpenAI and Perplexity pick up a modified robots.txt in about 24 hours. At Google, recrawling a modified page can take anywhere from several days to several months. The reading becomes meaningful when it is repeated monthly over a quarter.


Limits to keep in mind

  • Citation is not guaranteed. OpenAI says so for ChatGPT. Google says so for indexing and serving. No provider controls an assistant's final choice.
  • Answers vary. Wording, history, location and model version change the list. A single reading proves nothing.
  • An llms.txt file is not a Google lever. Google states that no AI-specific file is needed for its AI features. Other tools may read it. It replaces neither indexing nor published facts.
  • Measurement undercounts. Some visits from assistants arrive without a referrer and land in Direct in GA4. The figures you get are a floor.
  • An inflated figure backfires. The assistant repeats the published figure, and the buyer checks it in the interview.
  • GEO does not replace the sales relationship. A citation opens a first contact. The dossier, the conversation and the staffing remain the job.

Frequently asked questions

What is the difference between GEO and AEO?

AEO (Answer Engine Optimization) works on the shape of content so an engine reuses a passage as a direct answer: definition, short answer, table, FAQ. GEO (Generative Engine Optimization) targets the presence of a source in an answer a model writes from several pages, and the accuracy of what it says about that source. Both rely on sound SEO.

Should we block GPTBot to protect our content?

It is a training policy choice with no effect on visibility in ChatGPT Search. OpenAI documents two independent crawlers: GPTBot for model training and OAI-SearchBot for search. A firm can decline training and remain visible in answers, provided it does not block OAI-SearchBot. The same separation exists at Anthropic (ClaudeBot and Claude-SearchBot) and at Mistral (MistralAI-Training and MistralAI-Index).

Can a citation in ChatGPT or Perplexity be guaranteed?

No. OpenAI states that ChatGPT ranks results using multiple factors and that placement is not guaranteed. Allowing search crawlers makes a page eligible, nothing more. The right measure is a mention rate over a fixed set of questions, taken every month.

Is an llms.txt file necessary?

Not for Google, which states that no AI-specific file is required for AI Overviews and AI Mode. Other tools may read the file. It replaces neither search crawler access nor factual service pages, which remain the priority.

How can we tell whether ChatGPT sends visits to our site?

In GA4, the AI Assistant default channel has grouped visits from ChatGPT, Gemini, Copilot, DeepSeek and Grok since May 2026. ChatGPT also adds the utm_source=chatgpt.com parameter to links in its search results. For Perplexity, you need a custom channel group. These figures remain a floor, because some visits arrive without a referrer.

Where should we start with two weeks available?

Check robots.txt and the firewall, settle a single fact sheet (name, headcount, stacks, locations, lead time), rewrite the most requested service page with an In short block and an FAQ, and move ten dossiers to the structured template. Then ask ten client questions in ChatGPT and Perplexity and log the answers. The rest of the plan can follow the next month.


Structuring dossiers without turning it into an SEO project

A dossier generator does not place a firm inside ChatGPT. It addresses a frequent cause of poorly summarized dossiers: lack of time, which leads teams to resend an old Word document updated in a hurry.

StepWhat the tool handlesWhat the firm decides
TemplateThe same sections from one dossier to the next: role, skills, engagements, sectorsWhich three engagements to feature
Engagement write-upConsistent layout for context, role and deliverableWhere to draw the line between led and contributed
VocabularyThe same labels and levels across every dossierThe volumes and lead times actually delivered
ExportGenerated PDF with selectable text and a heading hierarchySending it to the client

The dossier that reaches the buyer is the one their assistant will summarize. If it is structured, the summary is more likely to be accurate.


Key takeaways

GEO is about being cited, and accurately, in a generated answer. AEO is about writing passages an engine can reuse on their own. Neither replaces SEO, and nobody can guarantee a citation.

Three workstreams for an IT services firm:

  • Access. robots.txt, firewall and HTML rendering let OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot and Bingbot through. Declining training remains a separate decision.
  • Facts. Every service page opens with a quantified In short block, answers buying questions in an FAQ and cites dated proof. The same figures appear on the website, in dossiers and in tender responses.
  • Measurement. Fifteen fixed questions, four assistants, one reading a month, supplemented by GA4, Bing Webmaster Tools and server logs. Inaccurate answers get fixed before absences.

An assistant repeats what the firm publishes. The work consists of publishing accurate facts that a machine and a buyer can both read.


Sources