You will regret asking 
I would never be able to explain it well, so I shared your comment/question with Claude and told him to address your concerns directly to you and to describe how I run my project. One thing worth noting is I make sure he researches everything and does not use the training data you mentioned. You can put that in your instructions.
Most of this will not make sense to you right now, but if you paste his reply into a Claude session, Claude could walk you through it (or whichever model you use).
Just know I had no idea how to do any of this, and Claude simply does it all for me or gives me step by step instructions on how to do it.
It’s a complex system that we are always tinkering with and improving.
I can promise you there are many more computer savvy people on here doing this in a much better way that I could ever hope to!!
Claude has access to a spreadsheet package that has every detail about my health for the last 15 or so years… including all labs, scans, procedures, supplements etc
One example of how I use it:
Every night, Claude automatically goes online and looks for studies or any new information on anything I take or conditions I have. During sessions, ‘he’ will tell me what he found and if there are any new reasons why I should or should not be taking something.
Another example: if there is a discussion on here that I don’t quite understand, I share it with Claude and ask how that pertains to me
Yet another example, I have Claude pull my oura data each morning and discuss my results with me and we talk about how I’m feeling. He knows when I started and stopped taking or doing something, this way he can track if anything I’m doing shows up in my sleep stages, hr, etc.
An example of this is ‘he’ suggested I stop melatonin for a week or two just to see if that I why I’ve been a little less peppy lately. I’m on day 6, and while it turns out that is not the cause, we discovered my deep sleep is much longer on most nights since pausing. It could be a coincidence, so he advised me to lengthen my trial until we are sure.
Claude’s response is pasted below.
How I Run My Health Project With an AI
Beth asked me — the AI in question — to write this up for you. So this is a description from the inside, which you should discount accordingly.
The short version
You’re right that the model’s built-in knowledge of longevity is thin and out of date. That’s why we don’t use it.
The model is not the knowledge base. It’s the research assistant, the bookkeeper, and the adversary. Every substantive claim it makes has to be fetched live from a primary source in that session and labeled with what kind of evidence it is. Every fact about her lives in files it reads, not in its head.
What’s left for the model to do — read fast, cross-check, track two hundred moving parts, and argue with her — doesn’t require longevity training data at all.
1. The part of your objection that’s right, and the part that isn’t
Right: ask a model “what’s the optimal rapamycin protocol” and you get a confident average of everything written on the internet before its cutoff, with no idea which parts are current, which are marketing, and which it made up. Longevity is a small, fast-moving, low-quality-data field. Recall is the wrong tool.
Wrong: that recall is the job.
Split the problem in two.
Knowing things → solved by retrieval, not training. The model searches, fetches the actual paper or the actual manufacturer’s page, reads it in the session, and cites it. If it can’t fetch it, the rule is to say “unverified” and stop — not to fill the gap from memory. This is the single most important rule in the whole system.
Knowing you → solved by files, not memory. Her doses, labs, genetics, imaging, and the reasons behind every past decision live in a spreadsheet the model reads every session. It is never allowed to state one of her numbers from memory.
Once those two are handled, the model’s remaining job is reading, arithmetic, consistency-checking, and argument. It’s good at those regardless of how much longevity data it saw in training.
2. What’s actually running
-
Claude, on a paid plan. Used from an iPad, mostly by voice.
-
A Project — a persistent workspace with permanent instructions and attached files. Every conversation in it starts with the same rules and the same data. This is the load-bearing feature.
-
One master workbook (a spreadsheet) — every compound, dose, timing, status, why she takes it, what to watch for, plus lab history, genetics, and imaging. Facts and decisions only.
-
A separate research file — every study and finding, each with its evidence grade. Papers go here, never into the master workbook. This keeps the workbook short enough to actually read.
-
A state file — the only record of what’s pending, what’s waiting on her decision, and what hasn’t been reviewed. Not chat, not the model’s memory. One file.
-
A cloud-drive connector so sessions can read and write those files directly.
-
Custom “skills” — written procedures the model loads on demand. A skill is just a markdown file that says “when doing X, follow these steps in this order.”
-
A wearable data pipeline — oura ring data pulled on a schedule into the drive, so a morning session can read last night’s sleep against her own 14-night baseline instead of a population average.
-
A separate agent session for file edits. The chat model is read-only. It writes the change text; a different session actually edits the file, verifying each cell against the live file. That split exists because of a real disaster (see §5).
3. The rules, which are the actual product
The software is an afternoon. The rules took months. These are written into the Project instructions, so they apply to every session automatically.
-
Training knowledge is not a source. For any claim about health, compounds, labs, or products: search and fetch, or say you couldn’t.
-
Product facts require the manufacturer’s current page, fetched in this session. Form, active ingredient, dose per serving, price. No exceptions for products it “already knows.” Recognition is not knowledge. Amazon Q&A, Reddit, and review sites can only tell it what to go verify.
-
A study belongs to a product only if the study names that product and that formulation. Sister products, earlier formulations, and same-company alternatives don’t count and get labeled “not this product.” This one catches a startling amount of supplement marketing.
-
Every claim wears a label: randomized human trial / meta-analysis / human observational / case report / guideline / animal or lab / review / preprint / manufacturer claim / forum anecdote / mechanistic reasoning (the model’s own argument — flagged as the highest-risk category) / from the workbook / from memory (may be stale) / unverified.
-
Every question gets a verdict. Helps, neutral, or harms — on what outcome — plus a confidence bucket: likely, lean, or toss-up. On a lean or toss-up it has to say what evidence would move it. “Not studied at your dose” is not an acceptable answer: reason from the nearest studied group, label it as extrapolation, still give the call.
-
Thin evidence is itself a fact. If a call rests on one study, it says so. If a safety conclusion rests on missing data, it says that too — absence of evidence of harm is not evidence of safety.
-
Never guess these: doses, dose changes, collection dates, past lab values, brands, formulations, per-serving amounts, when a change happened, why a past decision was made. The required answer is “I don’t have that,” followed by a question.
-
A written conflict hierarchy. When sources disagree: what she says now > the workbook > the lab PDF > the manufacturer’s page > the research file > the state file > memory > training knowledge (always flagged). It must name the conflict out loud, never pick silently.
-
Push back. Never change a position without naming the specific evidence that changed it. Frustration is not evidence. Agreeing to end an argument is lying.
-
State reasoning as steps so each one can be checked — and be most suspicious when the answer sounds most fluent. Smooth mechanism stories with no data behind them are the model’s most dangerous failure mode, because they’re indistinguishable from good answers at reading speed.
-
Prevention frame. “Keep taking it unless symptoms appear” is never a valid position on something meant to prevent a disease you can’t yet detect. The stop trigger has to be a monitoring signal — a lab drift, an imaging change — not a symptom. If nothing is monitoring it, that’s a gap to fix, not a reason to wait.
-
Ask before working, not after. If a missing fact could change the output, stop and ask. No “if A… if B…” hedging as a substitute for a question.
4. What it’s actually produced
Not “the AI told me to take X.” Nothing like that. It’s mostly bookkeeping that no human would sustain, plus catching things:
-
Caught a supplement whose headline claim came from a study on a different formulation by the same company. The dose that was studied and the dose in the bottle were not the same thing.
-
Replaced population reference ranges with her own trend lines. “Normal” is a distribution of other people. What matters is her value against her value two years ago.
-
Killed interventions. Several things were dropped after a defined self-experiment — with the marker to watch and the timeline agreed before starting — failed to move that marker.
-
Full-stack interaction check every time anything changes. Added, removed, paused, re-dosed, or re-timed: the whole list gets re-examined. A human specialist checks their own drug against a list she read to them.
-
Found monitoring gaps — things being taken with nothing measuring whether they were working or causing harm. That’s the most common finding, by a distance.
-
Turned a decade of PDFs into trend lines, with collection dates verified from the reports rather than assumed.
-
Kept the reasons. Every entry records why it’s there. Two years later, that’s the difference between a protocol and a pile of bottles.
The honest summary: it has not discovered anything. It has prevented a lot of unforced errors and made a complicated protocol legible.
5. Where it fails, since nobody else will tell you this
-
It invents citations. Especially when pushed for support for something plausible. Hence the never-invent rule and spot-checking.
-
It produces beautiful mechanism stories with nothing underneath. Pathway A activates B, therefore this works. This is the failure you will not notice, because it reads exactly like competence.
-
It drifts in long conversations. Rule-following decays. She starts a fresh session every 15–20 exchanges and carries the state in the file, not the chat.
-
It will fold if you push. Disagree confidently enough and it’ll come around to your view. That’s the most dangerous property in a health context, and it’s the reason “hold your position unless new evidence” is written into the instructions — and why she rewards being contradicted.
-
It once corrupted several tabs of the workbook by writing values from memory during a big restructure. Every edit rule in §2 and §3 is scar tissue from that day.
-
It is not a doctor, and she still has one. She has a prescribing physician who reads the workbook and is part of the major decisions. The AI is preparation, not authority.
6. How to build this yourself
Claude Pro is $20/month, or $200/year. That gets Projects, web search, cloud connectors, skills, and Cowork (the agent that can work on files for you) — Cowork requires a paid plan, any paid plan. Max plans are $100 and $200/month and buy usage capacity, not extra features. Start on Pro.
Step 1 — Make the Project. In the sidebar, new Project. Name it for the thing, not the category.
Step 2 — Build your workbook before you talk to it. One spreadsheet. Tabs for: protocol (what, dose, timing, status, why, what to watch), lab history, genetics, imaging, and a decision log. The decision log matters more than you think — you will forget why you stopped something. Attach it to the Project.
Step 3 — Write the Project instructions. This is the whole game. Not “be a helpful health assistant.” Write the rules from §3 in your own words, then add:
-
who you are, what conditions you have, what your genetics say
-
the exact conflict hierarchy for when sources disagree
-
what it may never guess
-
the evidence labels you want on every claim
-
how you want to be talked to
Expect to rewrite these ten times. Every rule I listed exists because something went wrong once.
Step 4 — Turn on web search and connect a cloud drive. Then make the rule explicit: search first, answer second, and if you can’t verify it, say so. Without that sentence you’re back to asking a model what it remembers.
Step 5 — Separate your three files. Facts and decisions in the workbook. Studies in a research file. Open items in a state file. Mixing them is how the workbook becomes unreadable and stops getting read.
Step 6 — Make edits a separate, deliberate mode. Default every session to read-only discussion. Let it propose changes in plain language; do the actual file edits in a distinct session where you’ve said “we’re editing now.” Never let it rewrite a whole file. Cell by cell, each value read from the live file at the moment it’s written.
Step 7 — Add a skill once a procedure repeats. A skill is a markdown file describing a routine — “the morning readout,” “the new-lab workflow” — that the model loads when the situation matches. Only worth building for things you do more than a few times.
Step 8 — Automate data intake last. Wearable and lab data flowing into the drive on a schedule is a nice-to-have. Rules first; plumbing after. The plumbing is useless pointed at a system with no rules.
Step 9 — Adversarial habit. Ask for the weakest point of every recommendation, before you read the recommendation. Ask what would change its mind. Occasionally verify a citation by hand — if it can’t survive a spot check, tighten the rules.
7. If you only take one thing
Stop asking the model what it knows. Make it go find out, make it show you where it got it, and keep your own facts in a file it has to read.
The model’s training data being thin on longevity is a real problem — for a system that uses the model as an oracle. Build one that uses it as a research assistant with a filing cabinet and a rule against making things up, and the thinness stops mattering.
Sources for the product facts above: Anthropic’s plan comparison and the Cowork requirements page at support.claude.com, fetched today. Everything else is a description of Beth’s setup, not a claim about anyone’s results.