Skip to content

The Boily engine

One measurement engine, from clinics to product pages

Boily for Commerce runs on the engine Boily built to measure how often AI assistants recommend clinics in Korea. This page explains what it measures, how it scores, and what it changes on your product page — and what it doesn't claim.

Measured on ChatGPT · Claude · Gemini · Perplexity

Measurement

The same test, every two weeks

A trend is only useful if the yardstick doesn't move. Everything in the measurement loop is fixed after the first round.

  1. 1

    Frozen buyer questions

    For each product and language, the first round writes the questions shoppers actually ask — needs, use cases, comparisons, brand questions. The set is saved and asked again, unchanged, in every later round.

  2. 2

    Pinned models

    Each assistant is locked to the model it used in the first round. A provider quietly swapping models shouldn't look like a change in your visibility.

  3. 3

    Four assistants, web search on

    ChatGPT, Claude, Gemini and Perplexity answer through their official APIs with web search enabled. We keep the sources each answer cites and, where the provider exposes them, the searches it ran.

  4. 4

    Names read by a model, not a keyword search

    A language model reads each answer and lists the products it recommends, in order. That ordered list is what gets scored, so a passing mention of a word isn't counted as a recommendation.

  5. 5

    Brand-gated matching

    A recommended name counts as your product only when it matches your product's core name and comes with your brand. Someone else's generic “vitamin C serum” is not you.

  6. 6

    Biweekly rounds

    Rounds run every two weeks on a fixed schedule. Each one is reported by assistant and by question, next to the previous round.

How the numbers are calculated

Two numbers per round: how often you appear (visibility), and how prominently (the recommendation score).

Named first
1.0
In the top three
0.7
Mentioned lower down
0.3
Not mentioned
0
Recommendation score per answer, averaged over every call in the round.

Errors stay in the denominator

A timeout or API error counts as a call where you didn't appear. Every round is divided by the same number of calls, so a bad API day can't inflate your score.

Partial rounds are labelled

If more than 20% of one assistant's calls fail, the round is marked partial instead of being shown as a normal result.

Text answers and shopping cards stay separate

Text answers come from the official APIs. Shopping cards in consumer apps aren't available there, so when we report them, it's as a separate, clearly labelled number.

Page improvements

What the engine writes, and the rules it follows

Between rounds, the engine writes improvements onto the product page from your own product data. Each one is checked before it's published, backed up, and can be rolled back.

Fact-first lead

A short opening sentence with your brand, the product's category and the two or three facts assistants look for. If a generated lead drops those facts, a plain template lead is used instead.

FAQ from missed questions

The exact wording of the frozen questions where at least half of the assistants didn't name you becomes the FAQ — up to eight, brand questions excluded.

Grounding checks

Every number, unit, ingredient, material and certification in the text must appear in your own product data. Price and rating come only from the product record.

No hedging

If your data can't answer a question, the question is left out. The engine never publishes “this information isn't available”-style filler.

Claims guard

83 rules covering Korean advertising, cosmetics, food and supplement and medical-device law, US FTC guidance and Japan's PMD Act and Premiums and Representations Act. Blocked sentences are removed; risky ones are rewritten once and checked again.

Server-rendered structured data

Written through each platform's server-side mechanism, because AI crawlers don't run JavaScript. It enriches the Product data your store already outputs rather than adding a duplicate.

Product feeds

Hosted feeds for Google Merchant Center, Microsoft Merchant Center and an OpenAI-format feed, enriched with the same lead and specs. You register them once.

IndexNow

After a change goes live, participating search engines are notified where the platform allows it. Where it doesn't, the report says so instead of counting it as done.

Brand profiles (sameAs)

Your brand's official profiles are linked in structured data — but only the ones you confirm. Profiles we detect automatically stay suggestions until you do.

Results vary. AI assistants are run by third parties and their answers change. The engine doesn't promise rankings or sales; it measures the same way every round and works on the parts of the page you control.

From clinics to commerce

The same engine, adapted to products

The engine started at Boily (boily.co.kr), which measures how often ChatGPT, Claude, Gemini and Perplexity recommend clinics in Korea when patients ask — with fixed question sets, biweekly rounds and the same rank-weighted score.

Boily for Commerce keeps that measurement core and adapts the rest to product pages: the unit is a product in one language, matching works on product and brand names, the claims guard covers consumer-product advertising rules, and improvements are written automatically through your store platform.

AspectBoily (clinics)Boily for Commerce
Unit measuredA clinic, for a specialty and areaA product, in one language
QuestionsPatient questions, frozen per setBuyer questions, frozen per product and language
Who appearsClinic namesProduct names, brand-gated
Content checksKorean medical advertising lawConsumer-product ad rules (KR · US · JP)
Page workDone for the clinic, on requestAutomatic, with backup and rollback

Results at clinics (Boily)

What the same engine recorded at clinics

First and latest rounds for some of the clinics Boily measures every two weeks. We didn't pick only the gains — clinics that fell or stayed level are here too. Names are masked down to area and specialty.

On the source page (data as of 2026-10-07), 19 of 23 measured sets have two or more rounds: 17 are above their first round and 2 are below it.

Visibility in AI answers (ChatGPT · Claude · Gemini · Perplexity)

  • ○○ Dental, Yeongtong (Suwon)

    Dentistry

    26.3% → 49.6%

    7/6 → 9/28, 2026 · 7 rounds

  • ○○ Dental, Ansan

    Dentistry

    3.6% → 31.7%

    7/25 → 10/2, 2026 · 6 rounds

  • ○○ Dental, Chungju

    Dentistry

    0.0% → 29.4%

    8/23 → 10/4, 2026 · 4 rounds

  • ○○ Hospital, Gwangju

    General hospital

    0.6% → 15.6%

    7/15 → 10/7, 2026 · 7 rounds

  • ○○ Dermatology, Geomdan (Incheon)

    Dermatology

    0.7% → 11.6%

    8/10 → 10/6, 2026 · 5 rounds

  • ○○ Dental, Gaepo (Seoul)

    Dentistry

    25.0% → 25.4%

    7/6 → 9/28, 2026 · 7 rounds

  • ○○ Dental, Magok (Seoul)

    Dentistry

    56.9% → 46.0%

    8/29 → 9/28, 2026 · 3 rounds

  • ○○ Dental, Changwon

    Dentistry

    25.4% → 21.7%

    9/17 → 10/1, 2026 · 2 rounds

Clinic results measured by Boily; product results vary and are not guaranteed.

  • Visibility is the share of answers that name the clinic, from a fixed question set per clinic asked to all four assistants through their APIs.
  • One assistant's measurement model changed on 19 August 2026, so rounds before and after that date aren't directly comparable (stated on the source page).
  • These are observations, not proof of cause. Each clinic was observed over a different period.

Source: boily.co.kr/research/cases

See the engine on your products

Pick your store platform to see how it plugs in, or open the sample report to see what a round looks like.