Skip to main content
Rewording a prompt barely changes which brands AI engines mention — until the rephrasing drifts past a measurable semantic threshold. Peec analyzed 1,754 prompts and 37,804 AI responses across five engines and 18 subverticals (published study, 2026; methodology in the SSRN working paper, Jun 2026), using two designs: 288 human-written prompts, and controlled variants stepped in tiny cosine-similarity intervals.
  • Human phrasing only looks chaotic. ~88–92% of human prompt pairs for the same commercial need sit above 0.50 cosine similarity; ~95% above 0.40. Less than 10% of variations drift semantically far.
  • Mentions hold above the threshold. As long as variants stay above ~0.50–0.60 similarity (engine-dependent), brand mentions stay stable. In the lowest bin (0.35–0.39), mention probability drops 2.40 percentage points against a 4.9% baseline — roughly a 50% relative decrease.
  • Middle-of-funnel is the sensitive zone. Unbranded commercial discovery prompts (“best CRMs for a small remote team”) drop off as early as the 0.60–0.65 bucket; top-of-funnel and branded bottom-of-funnel prompts are stable. Suggested tracking split: ~25% TOFU, 50% MOFU, 25% BOFU.
  • Beware the semantic blindspot: “car rental Munich” vs “car rental Frankfurt” are 95% similar but different intents — a changed qualifier (location, product, demographic, brand) is a new intent.
  • Engines differ: Google AI Overviews shows the most persistent MOFU sensitivity; Gemini’s effect fades fastest — report engines separately.
Takeaway (2026): prompt tracking works, but not as a one-prompt measurement system — track a portfolio across semantic neighborhoods, funnel stages, and engines. How prompt style shifts baselines: Prompt style and brand mentions. Source: How prompt wording impacts AI brand visibility, Peec via Search Engine Journal, 2026.