"The grader does not ask the internet what it thinks of your brand. It asks the LLM what the LLM thinks of your brand. That is a meta-evaluation, not a measurement."

When HubSpot launched its free AEO Grader in July 2026, the marketing positioned it as a tool that 'reveals what AI engines are saying about your brand.' That framing is accurate but incomplete. To understand what the scores actually mean, you need to know how the tool works under the hood. So we reverse-engineered it.

The short version: the HubSpot AEO Grader is not a web crawl. It does not simulate a real user query and observe what ChatGPT, Perplexity, or Gemini returns. It sends your brand information directly to each AI engine via API and asks the model to evaluate itself. The score you receive is the LLM's own assessment of how well it knows your brand. That is a meaningful distinction.

The Architecture

The grader takes four inputs: company name, geography, products or services, and industry. Once submitted, it fires three parallel API calls, one to each engine: OpenAI (specifically GPT-5.4 mini, not GPT-5 or GPT-4o), Perplexity, and Gemini. Each call has a 90-second timeout. The request payload for each call includes the four input fields plus a session hash, language code, and tracking source identifier.

Each engine call hits a scoring endpoint at aims-ai-grader.hubwt.com/v2/production/score, with a queue-based fallback at wtcfns.hubspot.com/wt-ai-grader-api/v2/score. A separate call to /v2/llm_reports retrieves the written interpretation of the scores. The PDF report is generated via a third service at pdf.hubwt.com. The full written report, brand archetype, confidence level, and improvement recommendations are gated behind a form submission.

HubSpot describes the scoring as 'deterministic' and uses structured output with JSON schema enforcement and retry logic to get consistent score formats from each LLM. The scoring rubric is applied by the model itself, not by a deterministic algorithm on HubSpot's servers. The model receives the brand context and the scoring criteria and returns a structured JSON object with numeric scores.

The Scoring Framework

The total score is 100 points distributed across five dimensions. Sentiment carries 40 points, the largest single weight. Presence Quality and Brand Recognition each carry 20 points. Share of Voice and Market Competition each carry 10 points.

Sentiment (40 points) covers three sub-dimensions: general sentiment (the overall positive, negative, or neutral tone of AI descriptions), contextual sentiment (how tone varies across topics such as customer support versus product innovation), and source-based sentiment (the credibility of sources influencing the AI's characterization). The 40% weighting reflects a structural reality: AI engines do not just recognize brands, they characterize them. A brand that appears in every AI answer about its category but is consistently described as 'complex' or 'best suited for large enterprises' will lose to a competitor described as 'intuitive' and 'fast to deploy.'

Presence Quality (20 points) measures mention depth, source quality, and data richness. Brand Recognition (20 points) measures how widely and specifically the AI can discuss the brand beyond surface acknowledgment. Share of Voice (10 points) measures the brand's rank relative to competitors in AI-generated responses. Market Competition (10 points) assesses how AI positions the brand relative to category peers.

What the Scores Actually Tell You

To calibrate the tool, we ran it on Salesforce. The results across the three engines were: OpenAI 87/100, Perplexity 85/100, Gemini 83/100. Brand Recognition was near-perfect across all three (19/20). Market Competition was perfect across all three (10/10). The variation appeared in Sentiment, where Gemini scored Salesforce 30/40 versus OpenAI's 33/40 and Perplexity's 34/40, and in Share of Voice, where Perplexity returned 5/10 versus 8/10 from both OpenAI and Gemini.

The Salesforce result illustrates the tool's primary use case: identifying which engine is characterizing your brand less favorably and in which dimension. A Sentiment gap between Gemini and OpenAI is actionable. It tells you that the sources shaping Gemini's understanding of your brand are producing a different characterization than those shaping OpenAI's. The next question, which the grader does not answer directly, is which sources those are.

Three Limitations Worth Understanding

The first limitation is the model choice. The grader uses GPT-5.4 mini, not GPT-5 or GPT-4o. This is a cost optimization. For well-known brands like Salesforce, the mini model has sufficient training data to produce reliable scores. For smaller or newer brands, the mini model may have less complete knowledge than the full model, which could produce scores that understate actual visibility in ChatGPT's consumer-facing product.

The second limitation is the meta-evaluation problem. The grader does not ask the internet what it thinks of your brand. It asks the LLM what the LLM thinks of your brand. The score reflects the model's self-reported confidence in its brand knowledge, filtered through a scoring rubric the model applies to itself. That is not the same as observing what the model actually returns when a real user asks a real question about your category.

The third limitation is the snapshot problem. The free grader produces a single point-in-time assessment. AI model knowledge changes as models are updated and fine-tuned. A score from July 2026 may not reflect the same model state as a score from October 2026. This is the primary commercial argument for HubSpot AEO, the $50/month monitoring product that tracks changes over time.

What It Is Good For

Despite those limitations, the grader is a genuinely useful diagnostic for three purposes. First, it provides a fast cross-engine comparison. Seeing that Gemini scores your brand 12 points lower than OpenAI on Sentiment is a signal worth investigating, even if the precise number is not empirically grounded. Second, it surfaces the dimension breakdown. Knowing that your Brand Recognition is strong but your Share of Voice is weak tells you something different than knowing your overall score is 72. Third, it is free and requires no account, which removes the cost barrier for a first AEO audit.

The tool is best understood as a structured prompt to three LLMs, not a measurement instrument. The output is the LLM's self-assessment of your brand's presence in its training data, organized into a consistent scoring framework. Used with that understanding, it is a reasonable starting point for an AEO audit. Used as a definitive measurement of AI visibility, it will mislead you.

References

[1] HubSpot, <a href="https://www.hubspot.com/aeo-grader" target="_blank" rel="noopener noreferrer" class="text-amber-700 underline">AEO Grader</a>, July 2026.

[2] HubSpot, <a href="https://www.hubspot.com/products/marketing/aeo" target="_blank" rel="noopener noreferrer" class="text-amber-700 underline">HubSpot AEO</a>, July 2026.

AEO Updates is published by The Prompt Group. Editorial decisions sit with the AEO Updates team, and any commercial relationship that touches a story is labelled on the page.