Eight choices appeared. One had to win.
300 AI-generated shoppers each saw eight nearby clinics, booked one, and explained why the other seven lost.
You need to book a medspa appointment. Eight nearby places are sitting on your phone. One is close. One has hundreds of reviews. One shows prices. Another looks beautiful but tells you almost nothing. Which one feels safe enough to choose?
I asked an AI model to make that choice 4,500 times. A better story barely changed the answer. The new clinic was chosen most often when price, ratings, review volume, consultation, safety, and aftercare all supported the same low-risk choice. That final bundle was never tested against improved competitors.
You can understand this page without reading Part I. That first study was simple: show an AI-generated shopper eight nearby clinics, make it choose one, and ask why the other seven lost. Do that 300 times.
This study asks the next question. What happens if the shopper sees a different version of the market? Make the new clinic easier to understand. Change its prices. Add stronger reviews and safety information. In another version, improve only the competitors. For each version, generate 750 choices and count how often the new clinic wins.
The answer did not rise neatly. A clearer story barely helped. Clear prices and stronger trust information did better. The final version combined both with stronger ratings, 220 additional reviews, warm consultations, and visible aftercare. It reached 64.4%. But that version was tested against the starting competitors, not the improved ones. The study found a compelling offer. It did not yet prove a moat.
You do not need to read Part I first. These are the three things from it that matter for the story ahead.
300 AI-generated shoppers each saw eight nearby clinics, booked one, and explained why the other seven lost.
The Ukrainian Village option was chosen roughly one out of every three times it appeared. The other three made-up locations were far behind.
When core prices were shown, a clinic was chosen about 24 times in every 100 appearances. With no visible price, it was chosen about twice.
The first study found a promising clinic concept. This study asks which separate changes make that concept more or less likely to be chosen.
Each simulated shopper chose one clinic and explained why the other seven lost.
Make prices clearer. Add reasons to trust. Let competitors improve. Count how often the new clinic wins.
What we see: the answer moves a lotThe first study asks, “What wins?” This one asks, “What changes make it stronger, and what changes make it weaker?”
Imagine the moment. You are thinking about Botox, filler, or a skin treatment. You search nearby. A handful of clinics appear. You scan distance, reviews, photos, prices, and anything that makes the place feel safe.
The first study tried to recreate that moment with AI. It generated 300 separate shoppers. Each shopper saw eight clinics, chose one, named two backups, and explained why the other seven lost. Four of the clinics were made-up versions of the same new medspa idea in different neighborhoods.
One neighborhood version did well. The other three barely registered. Clinics with visible core prices were chosen much more often than clinics with no prices. The written answers kept returning to the same concerns: reviews, trust, distance, and whether the place felt right.
That was useful, but it was only one view of the market. This second study asks the question again after changing what the shopper sees. The same 750 modeled shopper profiles appear in all six versions, using the same booking task. That lets us compare each profile's modeled answer across versions. It still does not turn those answers into real customer behavior.
Write down one version of the market. Ask the AI shopper which clinic it would choose. Repeat that choice 750 times. Then rewrite the market and do it again.
There are six versions. The first is the starting point. The second gives the new clinic a clearer story. The third changes its price menu. The fourth strengthens its ratings, reviews, and care information. The fifth leaves the new clinic alone and improves nearby competitors. The final version resets those competitors to the starting point and gives the new clinic the full bundle.
The original project contains all 4,500 answers, the 4,500 prompt records, the same 750 shopper profiles in every version, and the code that created each version. The short summary omitted the recipe. The underlying files did not.
There is still a catch. These are six separate “what if?” versions, not a clean ladder where only one detail changes at a time. The final version changes prices, ratings, review counts, wording, and a review excerpt together. We can describe the bundle and compare its result. We cannot give one ingredient all the credit.
Every version follows the same four-step rhythm.
For example: change the price menu, strengthen reviews, or improve only the nearby competitors.
An AI-generated shopper chooses between the new clinic and the nearby alternatives.
Each version gets the same number of simulated choices.
How often did the new clinic win? What reasons kept appearing?
The shopper is always choosing a clinic. What changes is the information around that choice.
Start with the plain version of the market. Then make the new clinic’s story clearer. Change its prices. Add stronger reviews and safety information. In a separate version, improve the competitors instead. End by giving the new clinic the full bundle while putting competitors back at the starting point. Click through slowly or play the sequence.
These versions are not cumulative steps. Each starts from its own copy of the market. They are six separate “what if?” scenes. That is why the final version does not tell us whether the full bundle can survive a competitor response.
Nothing special has been added yet. This is the starting version of the market.
In the starting version, the new clinic was chosen 97 times out of 750.
These are AI-generated choices, not results from real shoppers. The original study files retain each answer and prompt record; this public page shows the totals and selected examples.
If every improvement helped, the bars would rise from left to right. They do not.
The new clinic starts at 12.9%. A clearer story moves it to 13.5%, almost no change. Changed prices reach 31.7%. Stronger reviews and safety information reach 28.9%. In a separate version, only the nearby competitors improve and the new clinic falls to 11.3%. The final bundle reaches 64.4% against the starting competitors.
The underlying setup files explain the jump. In the final version, every new-clinic option received clear $12–14 pricing, a 0.2-star rating lift, 220 additional reviews, warm-consult and aftercare language, and a new five-star review tying all of those signals together. Nearby competitors kept their starting profiles.
The starting version won 12.9% of choices. The clearer-story version won 13.5%, only four extra choices out of 750.
The price version reached 31.7%. The trust version reached 28.9%. Both change several visible inputs, so neither isolates one cause.
In a separate version, nearby clinics improved their prices and booking experience while the new clinic stayed at its starting profile. It fell to 11.3%, or 85 choices.
It reached 64.4%, or 483 choices. But every ingredient moved together and the improved competitors were absent. This is a promising offer, not proof of a defensible advantage.
Every bar represents 750 AI-generated choices. A taller bar means the new clinic was chosen more often.
Nothing special has been added yet. This is the starting version of the market.
The new clinic explains more clearly who it is for and makes booking easier to understand.
The new clinic receives standardized $12–14 pricing and $0–50 consultation terms. Both the visibility and some actual price inputs change.
The new clinic gains a higher displayed rating, 160 more reviews, and stronger language about the provider, consultation, safety, and aftercare.
Only nearby competitors improve their pricing and booking information. The new clinic stays at its starting profile.
The new clinic gets clear $12–14 pricing, a 0.2-star rating lift, 220 more reviews, and one aligned story about the consult, safety, and aftercare. Competitors return to their starting profiles.
The bars tell us where the answer changed. They cannot tell us which single edit caused the change because some versions changed several things at once.
The final version did more than tell a better story. It made nearly every visible signal support the same low-risk choice.
The new clinic showed a $12–14 injectable price, a low consultation fee credited toward treatment, stronger star ratings, 220 more reviews at every modeled location, warm consultations, visible care proof, and aftercare. A new five-star review then said those exact pieces “all matched” and called the clinic “the safest easy choice.”
That makes the result useful, but not magical. The model was shown a clinic with clearer prices and materially stronger reputation evidence. In two locations, the advertised injectable price also became lower, so this was not merely a formatting change. The final version tells us the full bundle is compelling inside this simulation. It does not tell us how much credit belongs to any one ingredient.
These inputs come from the original code and choice-set files, not from an interpretation of the 64.4% result.
Neurotoxin at $12–14 per unit, a $0–50 consultation credited toward treatment, and visible package ranges for skin services.
This is the same pricing input used in the pricing-only version.
The starting rating was increased by 0.2 stars, rounded to one decimal and capped at 5.0.
The displayed final ratings match the trust-only version after rounding.
220 reviews were added to the starting review count.
This is 60 more added reviews than the trust-only version.
Transparent, high-trust neighborhood care with clear prices, warm consultations, and strong care proof.
This combines price clarity and trust language in one description.
The pricing, reviews, consult, and aftercare all matched; it felt like the safest easy choice.
A new five-star modeled review was placed first in the visible review excerpts.
The improved competitors from Version 5 were not carried into Version 6. The final result shows the new clinic at its strongest against competitors at their starting point. It does not show whether that advantage survives imitation.
The same 750 AI shopper profiles appear in the starting and final versions. The study compares each profile's two modeled answers and sorts them into four groups.
387 were labeled “another clinic at the start, new clinic at the end.” 96 had the new clinic in both versions. Only 1 went the other way. The remaining 266 had another clinic in both versions.
The big number is 387. Most of the reported difference between the first and final versions sits there.
Do not picture 387 real people changing medspas. No customers were followed over time. These are paired AI answers for the same modeled profiles, not observed switching.
Each dot is one paired result in the study summary. It is not a real customer journey.
Among the 483 choices of the new clinic, 354, or 73.3%, named Trust / Safety as the single top reason.
Across all 750 final-version choices, including choices of competitors, Trust / Safety was the top reason in 441. Reviews and brand aesthetic account for the 39 choices omitted from the earlier short summary.
But trust was not independently discovered after the run. The final input explicitly described the clinic as high-trust, added safety and aftercare language, and inserted a review calling it “the safest easy choice.” The result shows the model responding to a deliberately trust-heavy scenario. It does not prove what real shoppers value.
Among the AI-generated choices of the new clinic, Trust / Safety was the single top reason for 354 of 483.
The scenario itself was loaded with trust evidence, including “high-trust” language and a review calling the clinic “the safest easy choice.” This is evidence that the model responded to that stack, not independent validation of real customer demand.
“I want the option that makes the risk feel smaller before I care about the deal.”AI-generated shopper answer · trust first · new clinic chosen
“Clear pricing helps, but I still need enough proof that the consult will not feel rushed.”AI-generated shopper answer · value aware · new clinic chosen
“The closest place only wins if it also feels safe.”AI-generated shopper answer · convenience led · another clinic chosen
“I would rather pay a little more for the provider that explains the plan and the aftercare.”AI-generated shopper answer · review driven · new clinic chosen
The original results name one top reason for all 750 choices.
The 750 answers contained 4,169 reason appearances because one answer can mention several things.
Reasons for rejecting the new clinic appeared 508 times alongside 267 choices of another clinic. One answer can mention more than one reason.
The written answers also explain why another clinic won. The same objections keep coming back: too few reviews, unclear prices, too far away, not enough trust, or a poor fit for the treatment.
In the starting version, those objections appeared 1,281 times alongside 653 choices of another clinic. In the final version they appeared 508 times alongside 267 choices of another clinic.
The totals are larger than the number of choices because one answer can give several reasons. For example, a clinic can be too far away and have too few reviews. The important point is not the raw total. It is that deeper reviews remained the most common objection in both versions.
These bars show how often each objection appeared in the written answers. One answer can contain several objections.
The “new clinic” in this study is actually a group of four unnamed locations. The study calls them A, B, C, and D.
At the start, Location A accounted for 70 of the group’s 97 choices. Location D received only one. In the final version, the four locations received 177, 139, 105, and 62 choices.
This does not prove that four real locations should open. It tells us only that the final total was not being carried by one simulated location alone.
The bars count AI-generated choices in the starting and final versions.
The study also groups the choices by where the simulated shopper lives, what treatment they want, and whether they are new to medspas. The six groups below changed the most between the starting and final versions. They are useful clues for future interviews. They are not predictions about real people.
Every card compares the first and final versions. These are AI-generated groups, not estimates of real customer behavior.
116 picked the new clinic in the final version
124 picked the new clinic in the final version
163 picked the new clinic in the final version
97 picked the new clinic in the final version
95 picked the new clinic in the final version
42 picked the new clinic in the final version
Use these cards to decide whom to interview. Do not use them as targeting instructions or expected conversion rates.
Do not open a clinic because an AI-generated result reached 64.4%. Use the result to design three small tests with real people.
First, make price easier to understand. Publish common service ranges, explain what a consultation costs, and say what happens after someone clicks. Then measure completed bookings and completed visits, not just page views.
Second, make care easier to trust. Strengthen reviews, provider explanations, safety information, and aftercare details while keeping the offer and booking flow the same. If bookings do not improve, the trust idea gets weaker.
Third, test whether the result survives competition. Keep the full final bundle on the new clinic and give nearby competitors the clearer prices and easier booking from the competitor-response version. Then repeat the run. That is the missing round.
The simulation gives us questions. Real behavior has to answer them.
The study found a compelling offer. The missing round is whether that offer can defend itself.
No person booked a clinic. No patient was interviewed. No real shopper changed providers.
An AI model was repeatedly asked to choose among clinic options. The original project retains all 4,500 answers, all 4,500 prompt records, and the code that built the six versions. This page publishes totals and selected examples rather than every row.
The final recipe is known. The weakness is experimental: several ingredients changed together, the advertised price itself changed in some locations, and the final bundle was not tested against the improved competitors. The right use is to choose the next test, not forecast demand.
The first study put eight clinics on a phone and forced one choice. This study changed what appeared on that phone and asked the question again.
A clearer brand story barely moved the answer. Price changes and stronger trust evidence each did more. The largest result appeared only when price, rating, review volume, consultation, safety, and aftercare all told one coherent low-risk story.
So the takeaway is not “the new clinic will capture 64.4% of Chicago.” It is this: language alone is weak; aligned proof is powerful. The model chose the clinic when the visible evidence made “safe and easy” feel true rather than merely claimed.
But the study stopped one round early. The final bundle faced the starting competitors, not the improved ones. Part II found an offer worth testing. It did not show that competitors cannot copy it.
Distance gets you considered. Aligned proof gets you chosen. Defensibility is the next test.
Read the original story of 300 AI-generated shoppers choosing one clinic from eight nearby options.