- Human phrasing only looks chaotic. ~88–92% of human prompt pairs for the same commercial need sit above 0.50 cosine similarity; ~95% above 0.40. Less than 10% of variations drift semantically far.
- Mentions hold above the threshold. As long as variants stay above ~0.50–0.60 similarity (engine-dependent), brand mentions stay stable. In the lowest bin (0.35–0.39), mention probability drops 2.40 percentage points against a 4.9% baseline — roughly a 50% relative decrease.
- Middle-of-funnel is the sensitive zone. Unbranded commercial discovery prompts (“best CRMs for a small remote team”) drop off as early as the 0.60–0.65 bucket; top-of-funnel and branded bottom-of-funnel prompts are stable. Suggested tracking split: ~25% TOFU, 50% MOFU, 25% BOFU.
- Beware the semantic blindspot: “car rental Munich” vs “car rental Frankfurt” are 95% similar but different intents — a changed qualifier (location, product, demographic, brand) is a new intent.
- Engines differ: Google AI Overviews shows the most persistent MOFU sensitivity; Gemini’s effect fades fastest — report engines separately.
Rewording a prompt rarely changes AI brand visibility until similarity drops
Peec’s study of 37,804 AI responses (2026) — brand mentions hold while prompt variants stay above ~0.50–0.60 cosine similarity; below that they drop ~50%, hitting middle-of-funnel hardest.
Rewording a prompt barely changes which brands AI engines mention — until the rephrasing drifts past a measurable semantic threshold. Peec analyzed 1,754 prompts and 37,804 AI responses across five engines and 18 subverticals (published study, 2026; methodology in the SSRN working paper, Jun 2026), using two designs: 288 human-written prompts, and controlled variants stepped in tiny cosine-similarity intervals.
