The newest frontier LLMs now give feedback on scientific papers that 45 domain experts rated as good as — sometimes better than — the top human reviewer at Nature-family journals. If a chatbot can referee professional science, what can it do for a personal scientist?
This week I handed my raw genotype file to OpenAI’s new Rosalind Workbench and asked it one question about migraines. Twenty minutes later I had a bioRxiv-ready paper.
The latest LLMs are good enough to referee a Nature paper. I don’t find that terribly surprising but that’s the conclusion of a large expert study from CMU and KAIST. About 50 domain scientists rated almost 3000 individual criticisms — every point raised by human and AI reviewers of 82 Nature-family papers — for correctness, significance, and evidence. An OpenAI agent beat each paper’s top-rated human reviewer, 60.0% to 48.2%. Claude Opus 4.5 and Gemini 3.0 Pro were statistically indistinguishable from the best human. All three beat the worst human on every dimension.
The (somewhat minor) catch: LLMs raise more incorrect points than the top human (86% correct vs. 92%), they overlap heavily with each other (21% of AI–AI criticisms match, vs. 3% for human pairs, so three bots are not three opinions), and they’re weak on “subfield conventions” — the tacit rules a working scientist absorbs slowly on the job. (See PSWeek260508)
For professionals, this is a workflow question. For personal scientists it’s something bigger: the difference between having a reviewer and not having one. Nobody peer-reviews my magnesium experiment (PSWeek260514). Until LLMs, if I wanted a second opinion on my own data, the options were a social media query or a Quantified Self meetup. When I tested a doctors-only medical LLM against plain Claude (PSWeek260423), I found the general-purpose model just as good. Now you can get even better, general-purpose models packaged to fit a scientist’s workflow.
Rosalind
ChatGPT Rosalind Workbench is OpenAI’s new life-sciences workspace, in research preview inside the regular ChatGPT app. It runs on GPT-Rosalind, a model that combines frontier reasoning with “specialized tool orchestration across medicinal chemistry, genomics, wet-lab assistance, and other scientific applications.” In practice that means it already has the sequence viewers, alignment tools, and plotting you’d ordinarily have to install separately.
I tried it with the easiest prompt ever. I uploaded a 700,000-line text file — the raw output of a cheap consumer genotyping test like 23andme. (PSWeek250403) — and asked it to find markers that might reveal my risk of migraines. I had already ruled out EMF as a possible cause for my teenage daughter’s headaches (see PSWeek230706 ) so I wanted to know if I carry something that might run in the family.
After a few minutes of chugging, it gave me a table of four specific variants associated with migraine, and this summary:
I matched 85 of the 123 independent migraine loci from a large study of 102,084 cases and 771,257 controls. Thirty-five were exact lead variants; 50 were strong markers from the same loci. The weighted balance was roughly the 80th–85th percentile among the study’s European-ancestry reference population. Treat that as a research estimate—not an absolute lifetime-risk percentage or a clinical polygenic score.
Then I asked it to write the whole thing up in a format suitable for bioRxiv, the preprint server where professionals post work before peer review. Soon I had a professional-looking paper: detailed methods, a discussion section, useful charts, and citations that — I checked — were real. Read it yourself (PDF).
What the 80th percentile means
Okay fine, but how “scientific” are my results?
When I asked Claude to review it pointed out that the “large study” on migraine risk that Rosalind quoted is Hautakangas et al. 2022 in Nature Genetics, and it’s a good one: 123 loci, 86 of them new. But the same paper estimates that all the common variants together explain maybe 11% of migraine heritability. Rosalind’s score is built from the part of that 11% my genotyping chip happens to cover — 85 of 123 loci, and 50 of those 85 by proxy markers rather than the actual lead variants. Whether a proxy in the same neighborhood carries the same weight is somebody’s judgment call, and who knows how it was calculated. I wonder how much of my 80th percentile lives in those 50.
To its credit, Rosalind said so. “Research estimate, not a clinical polygenic score” is well put, and frankly it’s more honest than a lot of consumer genetics reports.
What about the paper it generated? Real citations, sensible methods, a discussion section that would please a graduate advisor — produced in twenty minutes from a question I typed on a Sunday afternoon. Back when I tested the first “AI scientist” tools (PSWeek250612), the citations were the weak link; now they check out. What hasn’t changed is that just because something fits the pattern of an academic paper doesn’t make it true. I’m my case, ChatGPT wrote me a well-formatted hypothesis about a sample of one, and the only real way to confirm it is the same as it’s always been: go to the real world and find real data. In other words, back to personal science.
Personal Science Weekly Readings
Speaking of peer review, it’s not like humans have been doing all that great of a job either.
A 2026 blinded study in a cardiology journal had an LLM review 40 real submissions alongside the human referees. The AI matched the final editorial decision 68% of the time; the humans, 67.5%. Neither number is impressive, which is kind of the real point here. (tldr; the bar was never that high.)
Ask Dan Ariely, one of the most famous behavioral scientists in the world. You may even have read his book, Predictably Irrational, that impressed all of us with some practical consequences of his idea that humans make bad decisions in predictable ways.
Well, the book that made him famous is based partly on his 2002 peer reviewed paper Ariely & Wertenbroch (2002) on deadlines and procrastination. It’s been cited thousands of times. But when Data Colada examined it more closely they found those spreadsheets were a little too spot-on to prove his thesis. Those spreadsheets — “last saved” by user “Dan Ariely”, had suspiciously convenient results. It’s Ariely’s second such finding; a 2012 honesty study was retracted in 2021. His response this month: after “more than two decades, and hundreds of experiments,” his memory is “insufficient.”
The charts below show (left) what you’d expect if the data had been left alone and (right) how it was presented in the paper. Somehow the grades were adjusted in a way that statistically gave a final result with the conclusion Ariely wanted. Peer reviewers, looking just at the means, didn’t see any shenanigans. But this type of manipulation is easy for an LLM to spot.

And in the “science-based policy that turns out to be wrong” department, Works in Progress makes the painstaking case that your careful recycling is probably a net harm: something like a quarter of what’s in your bin is contaminated and has to be washed — expensively — before it can be processed. Much (most?) of the time, landfill is cheaper and cleaner than the alternatives. (But if you read PSWeek250102, where somebody put an Airtag into a Starbucks recycling bin, you already knew that).
Finally, what all of this week’s readings have in common: the real checking was done by outsiders, not by the credentialed process that was supposed to catch it. That checking just got very cheap. Rosalind will write you a paper in twenty minutes; or it can tear apart a professional’s life work in ten. The gap between “professional” and “amateur” has never been narrower.
So what about my paper?
It’s easy to submit a paper to BioRXiv (I’ve done it before) but with this one I haven’t tried, at least not yet. Part of my reason is that, because I’m not a professional scientist, I feel no urgency to “publish”. Is this just me suffering “imposter syndrome“?
I believe that in a few years people will prefer AI-assisted writing over human writing, at least for non-fiction or scientific writing, where clarity and accuracy are critical. So what’s left for humans?
The answer is reputation. I feel responsibility for anything that goes out under my name and to be honest, I don’t feel like I understand “my” paper well enough to defend its conclusions. Rosalind, to me, was a fun exercise and I’m glad it’s now available to people like me, but the real consequence is not how much stuff you can publish, but whether other humans will care.
About Personal Science
Personal scientists test the claims that the rest of the world repeats — including the ones a very capable AI just wrote up for us. Nullius in verba — take no one’s word for it — was the motto of a society whose members could check each other’s work. The checking is cheaper now. That’s not an excuse to skip it.
We publish every Thursday. If you’ve tried Rosalind or one of the other AI research workbenches on your own data, let us know.




