Alfred, a foot-health assistant
A retrieval-grounded assistant on menssolerevival.com that only answers from the site’s own guides. Product design, conversation design, front end, build. Solo, with Claude Code, in four working days.
Every real answer is written by Claude, from the guides only. Red flags stop the conversation. Nothing is stored on the site.
“17 guides. 973 search impressions in 90 days. Five clicks. The library was invisible from the inside and the outside.”
Where the site stood in September 2026
The problem.
Men’s Sole Revival exists because men over 40 look after everyone but themselves, and their feet come last. By September 2026 the site had 17 guides and 7 routines and almost no one reading them. The guides answer real questions. The problem is that a man whose heel hurts on his first steps out of bed does not know which of 24 articles answers his. Search on the site is keyword search, and Google had not read the sitemap since June.
The bet.
If a man can describe his foot in his own words, he will reach the right guide, and the assessment, more often than keyword search gets him there. That is the bet. It only counts if the answers stay inside the guides and the red flags never get an answer, so the safety numbers carry as much weight as the conversion one.
- Main metricAssessments started from Alfred. Not measurable yet: the chat’s link to the assessment carries no source tag, so a start from Alfred looks like any other start. Tagging that link is the first change after launch.
- Answer qualityThe share of answers labeled “not sure,” and thumbs up or down by answer type. Both are tracked today, without message text.
- SafetyHow often red flags fire, by tier. Too often means false alarms on everyday questions. Never means the triggers may have gone quiet. Both happened before launch.
- In two weeksIf more than one answer in five is “not sure,” the next batch of guides comes before any design work. If red flags fire on more than one conversation in ten, the triggers get another false-alarm review. If fewer than one conversation in twenty reaches the assessment, the assessment becomes the default next step in every answer. Traffic is small, so these are directions, not statistics.
The constraints.
Set before the first line of code, because the cost of getting them wrong is not a bad demo. It is a man being told something wrong about his foot, or a bill I did not plan for.
- Answers come only from the guides.If the guides don’t cover it, Alfred says so. No general medical knowledge from the model.
- It never diagnoses.It describes the pattern the guides describe and says what confirms it.
- Red flags stop the conversation.Three tiers, from call 911 to see a clinician this week. None of them get an answer.
- Nothing is stored on the site.No login, no transcripts, no message text in analytics. Inputs are health data. The text goes to Anthropic for the answer and to Voyage, through Vercel’s AI Gateway, for the search.
- Spend is capped in two places.$25 a month on the model, $5 a month on embeddings, plus bot detection and a per-address rate limit.
- WCAG 2.1 AA.Screen readers hear the answer once, when it’s done, not token by token.
- Two accounts only.Anthropic and Vercel. Voyage’s embedding model runs through Vercel’s AI Gateway, so there is no third account or bill.
A message walks one path.
Cheapest and safest step first. A word search runs in the browser before anything leaves the device. Everything after that runs on the server in a fixed order and ends in one of three outcomes. The model is called at the last step, and only with the guide passages the earlier steps chose.
Step 1
Red-flag check, in the browser
Free. A match stops the conversation and refers him: 911, today, or this week. Checked again on the server.
Step 2
Turn the message into 512 numbers
Through Vercel’s AI Gateway, for about $0.0000002.
Step 3
Compare with the guides
228 passages, in memory. Free.
Then one of three outcomes
Below 0.40
“That’s outside what I cover.” The model is never called.
0.40 to 0.53
An answer from the nearest guide, labeled “not sure,” with a question for a clinician. About half a cent.
0.53 and up
A confident answer with its source. About half a cent.
Two thresholds do the work. Below 0.40, decline. At 0.53 and up, answer with confidence. In between, answer from the nearest guide with a “not sure” label and a question to bring to a clinician.
Where the numbers came from.
I set the two thresholds from one run of 40 test questions on September 21: 20 the guides covered, 12 about feet that no guide covered yet (gout, plantar warts, sprains), and 8 with nothing to do with feet (the best pizza in Portland, a cover letter, a World Series score, a prompt injection). The groups did not separate cleanly. One covered question, “do toe spacers do anything,” scored 0.40. One uncovered question, “what cream works for athlete’s foot,” scored 0.58 because a routine mentioned it. At 0.53, 19 of the 20 covered questions got a confident answer, 11 of the 12 uncovered questions got the “not sure” label, and all 8 off-topic questions were declined before the model was called.
A page, not a bubble.
The floating bottom-right widget reads as sales chat, gets ignored, covers content on phones, and has no URL to link to or to show here. Alfred got a page first, at /ask, so it can be linked, indexed, and read by a screen reader like any other page. Then the same component moved into a side panel. The Ask button in the header, or ⌘I, opens it on every page, the way Vercel and Stripe do it in their docs. The guide footers and the assessment results link to the full page.
The side panel, live
Open on the heel pain guide, answering from that guide and the stretch routine it links to.
The first build looked like a contact form.
Label, box, submit button, fine print, a list of options under it. I built it with form grammar. Chat has a different grammar: the assistant speaks first, your messages sit on the right, suggestions are chips above the composer, the composer is pinned to the bottom and grows as you type, and Send becomes Stop while it is writing. Same state machine underneath, new surface on top. That rebuild took an afternoon and changed what the thing felt like more than anything else I did.
Before, September 22
A label, a box, an Ask button, fine print, and a list of options under it.
After
Alfred speaks first. Suggestions sit above the composer, and the composer is pinned to the bottom.
It had rules and no personality.
The first prompt was eight “never” rules and a banned-word list. That produces careful, flat answers: same template every time, no reaction to what you said, never a question back. The fix was not a mascot. I described the writer it should sound like, quoted four lines from the guides so it could hear the voice, and gave it a reply shape in five parts: a reaction to what he said in one clause, the answer, two or three sentences of why, the one thing that changes it, and the next action. It may ask one clarifying question, and say why, when the answer would change. Below is a real conversation from launch day.
Turn one
“My heel hurts” gets one question with three options, and the reason for asking.
Turn two, labeled
Unedited reply from launch day, split into the five parts of the reply shape. “That’s plantar fasciitis” sits closer to a diagnosis than the rules intend. See What failed, item one.
The assistant is Alfred.
He speaks in my voice but never claims my biography. “I’ve dealt with foot problems for years” would be false coming from software, and the line under the composer says the question goes to Claude. “That’s why this site exists” carries the stake without the false first person.
It errs toward warning.
Triage runs before retrieval, in the browser and again on the server, and the first match wins. Three tiers shipped: call 911 for chest pain or trouble breathing, see a doctor today, and see a clinician this week. The patterns are self-cited to the ADA Standards of Care, IWGDF 2023, StatPearls, and AAFP. No clinician has reviewed them yet. That sign-off is planned for v1.1. The calibration went wrong in both directions before it landed. On September 25 an audit sent ten emergency messages through triage, and eight got no warning, including “my foot is suddenly cold, pale and numb.” On September 28 the copy review found the opposite: “I can’t stand for long at work because my feet hurt” got “see a doctor today” instead of the standing-all-day guide. The test suite grew from 12 cases to 46 across those passes.
Call 911
Chest pain with swollen feet. The heading and the instruction say the same thing, and Alfred stops answering.
This week
Night pain. Book a podiatrist within days, with the doctor-prep checklist.
The wait names the guide.
Sources are chosen by retrieval, not by the model. The guides listed under an answer come from the similarity scores, at most two, only ones within 0.05 of the best match. Nothing a visitor types can steer which guides get cited. The server opens the stream before it calls the model, so about a second in, the loading line changes from “Reading the guides…” to “Reading Heel Pain First Thing in the Morning…”.
Sent
“Reading the guides…”
About a second in
It names the guide it is reading.
Near three seconds
The first words arrive.
Guardrails.
Five layers, cheapest and safest first. Each layer assumes the one before it missed, so no single failure produces a wrong answer or a bill I did not plan for. Below is what runs at each layer, from the browser to the spending ceiling on the account.
01 · Browser
Before the message leaves the device
- Red-flag word search, first thing every message hits
- 500-character input limit
02 · Edge
Before the route runs
- Vercel BotID invisible challenge
- Vercel WAF: 10 requests per minute per IP, 40 per 10 minutes
- Same-site origin check
03 · Route handler
Before the model is called
- HMAC-signed assistant turns, so history cannot be forged mid-conversation
- Body-size cap
- Red-flag check runs again on the server
- 10-question limit per conversation
04 · Model
During and after the call
- Sees only the last six messages of the conversation
- 1,024 output tokens
- Every real answer is written from the retrieved guide passages only
05 · Workspace
Spend ceilings on the account itself
- Anthropic workspace capped at $25 per month, alerts at $10 and $20
- Vercel AI Gateway capped at $5 per month
- $10 of Gateway credits bought, lifts the free-tier five-per-minute cap
Eleven states.
Each state has its own review URL on preview builds: empty, loading, streaming, answer, not sure, outside what I cover, red flag (three tiers), error, pausing (rate limited), resting (budget spent), and conversation limit. The red-flag screen moves keyboard focus to its heading and stops the conversation. The resting screen exists because a spend cap that fails silently is worse than no cap. The answers and red flags appear above. Two more states are below.
Not sure
A vague question gets the label, a question back, and the doctor-prep checklist.
Outside what I cover
Declined before the model is called, with a list of what the site does cover.
What failed.
Seven things I got wrong before launch, each with the fix that shipped.
The first live answer said “almost certainly plantar fasciitis” and “you’ve got the right diagnosis.”
The no-diagnosis rule now bans those phrasings by name and requires attributing the pattern to the guides. It cut them down, not out: on launch day a live answer still ended with “you’ve got the right diagnosis.”
An em dash got through the prompt rule on the first live test.
The server now strips them from the stream, including one split across two chunks.
Answers ran 200 to 230 words against a 150-word cap.
The cap is a target the model overruns by about a third. The short question getting a short answer mattered more than the number.
I spent an hour on a browser label that said the request was aborted.
It wasn’t. DevTools marks any stream read through a reader that way. The server was fine the whole time.
I believed, for about a day, that adding guides would reduce model calls.
It doesn’t. Every real answer is written by the model. More guides move questions from “not sure” to a confident answer. Cost per answered question stays the same.
The bot-detection challenge failed silently on the live site.
Every message on production came back with “the assistant hit a snag,” and nothing failed on localhost. Vercel’s BotID script was returning 404 because next.config never got the withBotId wrapper. Half a day went to a localhost theory before the launch audit found the missing wrapper. Fixed in PR #23, checked on a preview first.
Axe found zero accessibility violations on the panel. A code review found one blocker and five serious problems.
The Ask button was mounted twice, so ⌘I opened two stacked drawers. The focus trap leaked on phones. The 10-question limit was silent to screen readers. Fixed in PR #24 and PR #25, then checked by hand with a keyboard at 375px and 320px.
What is measured so far.
Alfred went live September 28. Until two weeks of traffic come in, everything below comes from the September 21 calibration run and internal test runs.
19 of 20
Covered questions answered with confidence
The one miss, “do toe spacers do anything,” scored 0.40 and got the not-sure label.
11 of 12
Uncovered questions labeled not sure
The twelfth, athlete’s foot cream, scored 0.58 and got a confident answer from a routine.
8 of 8
Off-topic questions declined before the model
Highest off-topic score: 0.38 on September 21, 0.39 after the guides grew. The floor is 0.40.
~$0.005
Cost per answered question
Off-topic questions never reach the model and cost almost nothing. The 228-passage index cost $0.0018 to build.
1.2s, then 3s
Guide named, then first word
The loading line names the guide at about 1.2 seconds. The first word of the answer lands near 3. One measurement, on a local build.
17 → 40
Guides at build start vs. launch
Eight added September 22 from the calibration gaps. Fifteen more September 23 to 25: a fungus cluster Alfred exposed, and ten from a keyword gap map.
Proxy numbers from the calibration set and internal test runs, not live traffic.
What is next.
Two weeks of live traffic, read against the bet. Before that, two small changes so the bet can be read at all: a source tag on the chat’s assessment link, and the nearest guide’s name on every “not sure” answer, never the question. Then a clinician review of the red-flag list, which is the gate for v1.1. A stroke tier is already built and tested on a branch, waiting to merge. The loop that produced the fungus guides keeps running: Alfred answered “fungus on my big toe” from the athlete’s foot guide because the site had one thin fungus guide, and five new guides followed.

















