Can Residency Programs Detect ChatGPT? ERAS AI Rules
Can residency programs detect ChatGPT? Studies say usually not — but ERAS makes you certify your own work. The real AAMC AI rules, and why detectors fail.
Can Residency Programs Detect ChatGPT in Your ERAS Personal Statement?
Short answer: mostly, they can't reliably tell. In every published study to date, program directors misjudged AI-written personal statements more often than they caught them, and no part of ERAS runs an automated AI scan on your essay. But "will they catch the AI?" is the wrong question, and the anxiety behind it pushes you toward exactly the writing that loses interviews. When you submit through MyERAS you certify that the statement is your own work; the AAMC permits AI only for brainstorming, proofreading, and editing; and the filter a good reader actually applies catches generic prose whether or not anyone labels it "AI." If you understand how AI detectors actually work and where they fail, you will see why "make it undetectable" is a trap — and why a structured residency personal-statement review aims you at ownership instead.
The rest of this piece gives you the real rules, the actual evidence, and the one test that decides whether your statement works.
The short version — detection is unreliable, and it is not the real risk
Two things are true at once. First, program directors are demonstrably bad at spotting AI prose in personal statements, and ERAS does not auto-scan submissions for it. Second, none of that helps you, because a statement that reads like it could belong to any applicant fails regardless of who or what wrote it.
So reframe the problem. The risk is not a scanner catching you. The risk is a certification you signed and a reader who is unmoved. Solve for those two, and "detection" stops mattering.
What ERAS and the AAMC actually allow (the real rules)
The AAMC's AI policy: brainstorm, proofread, and edit — the final work must be yours
The AAMC's position on AI in its application services is narrow and clear: you may use AI tools to brainstorm ideas, proofread, and edit, but the submitted personal statement must remain your own work. That is the whole rule. There is no ban on touching an AI tool, and there is no permission to have one write your statement for you.
Note the sourcing caveat: the AAMC's MyERAS guidance was browser-verified on 2026-07-20 (AAMC pages block automated fetchers), so treat the wording as confirmed as of that date rather than freshly machine-checked at this writing. The substance has been stable across the cycle: AI as an assistant, not an author.
What you certify when you submit MyERAS (attestation, not a scanner)
Here is the part most "will they catch me" articles skip. ERAS enforcement runs on attestation, not detection. When you certify and submit in MyERAS, you affirm that the material — the personal statement included — is truthful and your own. There is no automated AI checker sitting between your draft and the program's inbox.
That changes the nature of the exposure. It is not a technical-detection problem where the winning move is a better disguise; it is an integrity commitment you made in writing. Framed honestly: ERAS will not flag your statement as "AI," but you signed that it is yours. A statement you did not actually write is a certification problem no humanizer can launder away.
The 28,000-character ceiling is not a target
While we are on the rules: MyERAS caps the personal statement at 28,000 characters, but that ceiling is a system limit, not a goal. Readers still expect a one-page statement — see the 28,000-character ceiling versus the one-page reader norm for the length math. This matters here because padding a thin, AI-drafted statement toward the ceiling is a reliable way to make it read more generic, not more complete.
No, the NRMP is not "integrity screening" you — that memo is fake
If your anxiety traces back to a screenshot claiming the NRMP is scanning applications for AI — and that international medical graduates are being singled out — you can put it down. It is fake.
In September 2024, the NRMP publicly stated that a memo circulating on social media, which claimed the NRMP was conducting "integrity screening" of applications for AI use, was fabricated. In the NRMP's own words, from its newsroom response published September 13, 2024: "This statement is fake and was not issued by the NRMP." The NRMP went further: "The NRMP has not taken any stance on the use of AI-generated content or 'integrity screening', nor do we provide guidance to programs on screening candidates." The organization added that the fake statement does not reflect its values or mission and reaffirmed its commitment to equitable Match access for IMG applicants.
As of this writing, the NRMP has issued no AI policy and provides no guidance to programs on screening applicants for AI. The IMG-targeted version of the fake memo is the most harmful strain precisely because it preys on the applicants with the most at stake in the Match — so if a forum thread or advisor has you convinced the NRMP is hunting for AI in your file, the authoritative source says otherwise. There is no such program.
What the research actually shows programs can (and can't) detect
Program directors usually can't tell
The published evidence is consistent in direction, so read each study in its own terms rather than as one blended number:
| Study | Specialty / year | What reviewers found |
|---|---|---|
| Johnstone, Neely, Sizemore — Journal of Clinical Anesthesia (2023) | Anesthesiology, 31 program directors | 19 of 31 (61%) found nothing to distinguish the AI "athletic experience" statement from an applicant's own; 24 of 31 (80%) for the "cooking experience" version. 28 of 31 (90%) rated the AI statement acceptable; 22 of 31 (74%) rated it good or excellent. |
| Chen, Tao, Park, Bovill — "Can ChatGPT Fool the Match?" Plastic Surgery (2024) | Plastic surgery, 2 retired surgeons, 22 statements (11 AI / 11 human) | Overall accuracy telling AI from human = 65.9% (72.7% and 59.1% for the two reviewers; Cohen's κ 0.374, "fair"). No significant score difference between the AI and human statements (P = .4129). |
| Menon, Solomon, Berenson, Kushnir, Shapiro — Cureus (2025) | Otolaryngology, 8 blinded evaluators, 10 statements (5 AI / 5 human) | No statistically significant differences in readability, originality, persuasiveness, or interview-desirability (readability: AI 3.63 vs human 3.53). One otolaryngologist scored 9/10; one attorney scored 50%. |
A note on a number you will see quoted online. The "61–80% of program directors couldn't tell" figure is the anesthesiology study specifically — its two prompts scored 61% (athletic) and 80% (cooking). That is a within-study range, not a cross-study meta-statistic. The plastic-surgery study (reviewers wrong roughly a third of the time, with no scoring gap between AI and human essays) and the otolaryngology study (no statistically significant difference on any axis) point the same way without collapsing into a single invented percentage.
The takeaway is not "so you can get away with it." It is that human readers cannot reliably separate competent AI prose from competent human prose — which means a fluent, on-topic, utterly forgettable statement sails through the AI question and dies on the interview question anyway.
AI detectors are unvalidated and dangerous to honest applicants
The counterpart study is the one every program considering a detector should read. In Cureus (2025), Cumbo, Williams, Canterino, Aikman, and Baum asked whether residency programs could detect AI use with off-the-shelf tools — and found the tools are not fit for the job. GPTZero scored roughly 92–93% accuracy on known AI text but flagged 18–91% of real 2023 applicant statements as potentially AI. Winston AI called 0% of verified AI samples human, yet rated genuine applicant statements anywhere from 3% to 100% human. Undetectable AI called nearly all of the real and older statements 100% human. The authors' conclusion is the line to remember: "the use of invalidated tools may harm honest applicants."
The otolaryngology study found the same fragility from a different detector — Scribbr's tool flagged 93.6% of ChatGPT text but still produced false positives on roughly 4% of genuine human writing on average. This is the same false-positive problem that plagues AI detectors across college admissions: a detector score is a probability, not a verdict, and a wide false-positive band means an honest applicant can be flagged for prose they wrote themselves. A responsible program cannot use these tools to make consequential decisions — and the ones with real AI guidance say so. The University of Washington's GME AI guidance, updated in December 2025, explicitly warns that "Many AI detectors produce false positives or negatives, making them an imperfect tool," and asks for authentic voice and transparency rather than a detector gauntlet.
The test that actually matters — the paste test
If detection is unreliable in both directions, what does separate a strong statement from a weak one? A good reader is not running a classifier in their head. They are applying something much older and much harder to fake.
The AMA put it plainly. Sanjay Desai, MD, the AMA's Chief Academic Officer, framed the operative test for GME personal statements (August 27, 2025): "If we can cut and paste your paragraph into somebody else's personal statement, it's not personal enough." His colleague John Andrews, MD, the AMA's VP of GME Innovations, added the mechanism: "To the degree that you use AI, it distances it from the personal."
Run your draft through it. If a paragraph could be lifted into any co-applicant's statement without breaking — swap the name, nobody notices — it fails the paste test. And here is the connection to everything above: generic AI output fails the paste test by construction. A language model produces the statistically likely sentence, which is precisely the sentence that could belong to anyone. It has no access to the specific patient, the specific decision, the specific thing you got wrong and how you knew. The paste test is not an AI detector. It is a specificity detector — and it catches hollow writing no matter who typed it.
Why "humanizing" tools don't solve your problem
This is where a whole industry is selling anxious applicants the wrong fix. AI humanizers and detection-bypass tools promise to rewrite AI text so it "passes" as human. Set aside that the detectors themselves are unreliable, so the metric these tools optimize is noise. The deeper problem is a category error.
A humanizer changes surface wording. It cannot add what the paste test demands: the falsifiable, applicant-owned detail — the scene only you were in, the reasoning only you did, the outcome only you can account for. Rephrasing a generic paragraph produces a differently generic paragraph. You can move a detector's needle and not move a single reader, because the interview decision was never about lexical texture; it was about whether the person on the page is specifically you.
So a humanizer optimizes the metric (a detector score) while leaving the goal (an interview) untouched. That is the entire trap in one sentence. Chasing "undetectable" spends effort on the one variable that does not decide your Match.
The honest workflow — use AI to think, not to write your story
None of this is an argument against ever opening an AI tool. The AAMC-permitted uses are genuinely useful, and used well they make your statement more yours, not less:
- Brainstorm and pressure-test angles. Ask it what a program director in your specialty tends to have seen a thousand times, so you can avoid the cliché rather than write into it.
- Outline and sequence. Get help ordering your material — but supply the material yourself.
- Proofread and tighten. Grammar, transitions, a sentence that runs long. This is editing, and it is allowed.
The line is ownership. The story, the specific scenes, and the reflection have to be yours — because those are the only parts that survive the paste test and the only parts a reader remembers. That is also what the people reading your file tell us they weigh: see what program directors say they actually read and the criteria a residency reader applies. Notice that "did they use AI" is not on either list. "Did this person show me something specific and true" is.
That is exactly what a good review checks for. A structured, low-drift review reads your draft against those reader-side criteria and tells you where it goes generic — which paragraphs fail the paste test, where the reflection is asserted instead of shown, where you named a specialty without earning it. It does not rewrite your essay, and it makes no promise to help you "pass a detector," because that promise is worthless. The point is to make the statement unmistakably yours.
Bottom line
Can residency programs detect ChatGPT in your ERAS personal statement? Reliably, no — human readers miss AI prose more than they catch it, no ERAS scanner checks for it, and the "NRMP integrity screening" memo is fake. But that is the wrong thing to optimize. You certified the statement is your own work, and the reader's real test is whether any of it could belong to someone else. Humanizer tools cannot pass that test; only specific, owned experience can. Use AI to think and to tidy, write the story yourself, and put your energy into detail no model could have invented.
When you are ready for feedback, our medical-school essay review works the same way across the application landscape, and a dedicated residency personal-statement review will tell you where your draft still reads like anyone's.
Sources
- Johnstone RE, Neely G, Sizemore DC. "Artificial intelligence software can generate residency application personal statements that program directors find acceptable and difficult to distinguish from applicant compositions." Journal of Clinical Anesthesia, 2023. PubMed 37336139; ScienceDirect S0952818023001356.
- Chen H, Tao B, Park B, Bovill E. "Can ChatGPT Fool the Match? Artificial Intelligence Personal Statements for Plastic Surgery Residency Applications: A Comparative Study." Plastic Surgery, 2024. PMC11561920; DOI 10.1177/22925503241264832; PubMed 39553535.
- Cumbo N, Williams A, Canterino J, Aikman R, Baum J. "Can Residency Programs Detect Artificial Intelligence Use in Personal Statements?" Cureus, 2025. PMC12392696; cureus.com/articles/370285.
- Menon S, Solomon M, Berenson A, Kushnir L, Shapiro J. "Distinguishing Between AI-Generated and Human-Written ERAS Personal Statements in Otolaryngology." Cureus, 2025. PMC12799199.
- NRMP. "Response to the Fake Statement on the Use of AI and Application Screening." NRMP newsroom, September 13, 2024. nrmp.org/about/news/2024/09/nrmp-response-to-the-fake-statement-on-integrity-screening-for-residency-applications/.
- AMA. "For GME personal statements, rely on emotional intelligence not AI." August 27, 2025. ama-assn.org/medical-students/preparing-residency/gme-personal-statements-rely-emotional-intelligence-not-ai.
- AAMC. MyERAS Personal Statement guidance and "Use of Artificial Intelligence in AAMC Service Programs" (AI permitted for brainstorming, proofreading, and editing; final work must be the applicant's own). Browser-verified 2026-07-20. students-residents.aamc.org/applying-residencies-eras; aamc.org/services/use-artificial-intelligence-aamc-service-programs.
- University of Washington GME. "AI guidelines for residency and fellowship applications." Updated December 2025. sites.uw.edu/uwgme/ai/.
Review Your ERAS Personal Statement
Check specialty motivation, clinical evidence, reflection, and readiness.