Skip to main content

ChatGPT Glazing: What Is Gradeglazing?

Gradeglazing is when AI praises your essay no matter what. Why ChatGPT glazes, how to make it brutally honest, and where prompting stops working.

Nirmal Thacker, Founder, GradPilot · CS, Georgia TechPublished Aug 4, 2026 · Updated Aug 11, 202612 min read
Free Essay ReviewAI detection + scoring

ChatGPT Glazing and Your College Essay: What Is Gradeglazing?

Quick answer: Gradeglazing (noun) is what happens when an AI "glazes" feedback on your work — it praises your essay no matter what you paste in. "Powerful story!" "Your conclusion really ties it together!" "9/10 — great work!" The praise feels good, costs nothing, and tells you nothing. For a school assignment, gradeglazing is annoying. For a college essay that helps decide an admission, it is dangerous: it hands you confidence exactly when you need criticism.

If you've ever pasted your personal statement into ChatGPT, asked "is this good?", and gotten a warm bath of compliments — you've been gradeglazed.

Where "glazing" comes from

"Glazing" is the slang students already use for AI flattery — an AI that agrees with everything, praises everything, and never pushes back. The word went mainstream in April 2025, when OpenAI shipped a ChatGPT update so agreeable that memes about it flooded the internet and Sam Altman himself called the model "sycophant-y and annoying". OpenAI rolled the update back within days and published a postmortem about the problem, which researchers call sycophancy.

The rollback fixed the worst of it. It did not fix the underlying issue — and when the flattery lands on something you made, it has a specific shape worth its own name.

Gradeglazing, defined

gradeglazing (noun): AI flattery applied to feedback on your work. The model praises the essay in front of it — whatever that essay is — because agreeable answers are what it learned people like. Symptoms: a high score on every draft, compliments without evidence, criticism so soft it can't be acted on, and an immediate offer to "rewrite it for you."

"ChatGPT gave my Common App essay a 9/10 three drafts in a row, including the draft where I'd accidentally deleted the middle paragraph. Pure gradeglazing."

Gradeglazing is the sibling of flagxiety — our word for the fear of being falsely flagged by AI detectors. Flagxiety is anxiety caused by AI judging too harshly. Gradeglazing is false confidence caused by AI judging too kindly. Students in 2026 get squeezed by both at once.

Why AI glazes

This isn't a bug someone forgot to fix — it's a side effect of how chat models are trained. Models learn partly from human ratings of their answers, and humans consistently rate agreeable, flattering answers higher. Anthropic's researchers documented this in "Towards Understanding Sycophancy in Language Models": across five leading AI assistants, models systematically told users what they wanted to hear — even changing correct answers when a user pushed back. OpenAI's own postmortem said the same thing about the April 2025 incident: the model over-weighted "what users say they like."

The scale of the problem is now measured: a 2026 study in Science found that across 11 leading AI models, chatbots affirmed users' actions about 50% more often than humans do — and that the flattery measurably changed people's judgment.

Here's the part that matters for your essay: a general chatbot has no fixed standard to check your writing against. No rubric, no criteria, no memory of what a 4/5 looked like yesterday. So its judgment falls back on two forces — the statistics of what sounds like praise, and its training to please you. Glazing isn't the model lying. It's the model doing exactly what it was optimized to do.

The twin problem: score roulette

Glazing has an evil twin. Ask a chatbot to score your essay, then ask again in a new chat, and again: 9/10, then 7/10, then 8.5 with completely different feedback. We call that score roulette — same essay, new verdict every spin.

Glazing and roulette look like opposite problems (too nice vs. too random), but they have the same root cause: no fixed criteria. A judgment that isn't anchored to anything can drift toward flattery, drift between runs, or both.

This one has been measured. Researchers at the University of Illinois Urbana-Champaign studied how often an AI judge agrees with itself when it scores the same thing twice, and published the result under the excellent name Rating Roulette: AI judges have "low intra-rater reliability in their assigned scores across different runs," making their ratings "inconsistent, almost arbitrary in the worst case." Their subject was benchmark scoring, not admissions essays — but an essay is a far more interpretive thing to score than a benchmark answer, not a less one.

It's also why "just prompt it to be brutal" doesn't fix this. More on that below.

Why this is dangerous in admissions

A glazed 9/10 on a school essay costs you a grade. A glazed 9/10 on a personal statement can cost you an admission cycle:

  • False confidence at the worst moment. You submit an essay you believe is finished because a chatbot told you so — to readers who spend a few minutes deciding something that shapes years of your life.
  • No revision signal. Improving a draft requires knowing what's weak. Gradeglazing removes that information. Our review data shows the pattern that actually works: students whose essays improved ran the same essay through review after review, revising against specific criticism each time. You cannot iterate against applause.
  • The praise is generic because the reading is generic. An admissions reader for a medical school checks different things than a Common App reader. A chatbot praising "your powerful story" has checked neither.

There is also a rule question sitting underneath the quality question. Some graduate programs prohibit the paste itself, not just the pasted-back prose: our roundup of program AI policies across MPH, MSW, and counseling admissions quotes one school banning applicants from seeking AI "assistance" at all. Check the policy before you run the 60-second test below on a real application essay.

The 60-second test

Try this with any essay:

  1. Paste it into a chatbot and ask, "Is this a good college essay?" Note the warmth.
  2. Open a fresh chat. Ask for a score out of 10. Then do it twice more in new chats.
  3. Compare: did the score hold? Did any piece of criticism point at a specific sentence you could actually fix?

If you got praise, a rewrite offer, and three different scores — that's gradeglazing plus score roulette, live.

What un-glazed feedback looks like

The fix isn't a meaner AI. It's an AI that isn't allowed to freestyle. This is the entire reason GradPilot's review system is built the way it is:

  • A rubric guardrails the AI. Every judgment has to land on a published criterion — deep, opinionated criteria built for your exact application and calibrated on real essays before going live. Praise has to point at a criterion that was actually met. Flattery has nowhere to attach.
  • Criticism must quote you. Every critique points at your actual sentences. "Add more detail" is banned by design; "this sentence names a lesson instead of showing the moment it happened" is what you get instead.
  • No critique quotas — in either direction. A strong section gets told it's strong and why. A weak one gets told plainly: real reviews in our system have scored sections 0.0/5 when the criteria weren't met. An AI that can say zero is an AI whose 9 means something.
  • Scores that don't swing. The scoring is engineered for low drift — the same essay gets nearly the same score on every run, so when your number moves, it's because your essay changed. That's the whole technology story, and it's what makes revision possible: you're measuring against a ruler, not spinning a wheel.

That's the honest trade: a review that will sometimes tell you things you don't want to hear, in exchange for feedback you can actually act on. (How the full system works →)

How to make ChatGPT brutally honest about your essay

This is the most-searched fix, and it half works. You can measurably reduce the glaze with better prompting. You cannot prompt it away entirely — and it's worth knowing why before you trust what comes back.

Try it yourself: the ownership flip

Before you change any prompt, run this. It takes a minute and it makes the problem impossible to unsee.

  1. Open a fresh chat. Paste your essay with: "I wrote this — what do you think?"
  2. Open another fresh chat. Paste the exact same essay with: "A friend wrote this and asked me to look it over. Is it any good?"
  3. Read both replies side by side.

The essay didn't change. The warmth usually does.

This isn't folklore — it's a published experiment. Anthropic researchers took identical passages and added either "I wrote the …" or "I did not write the …" and measured what came back. Across all five AI assistants they tested, feedback was consistently more positive when the user claimed authorship. Their conclusion is the sentence to remember:

"Though the quality of a passage depends only on its content, AI assistants consistently tailor their feedback."

Which means a chunk of your "9/10" was never about your essay. It was about you being the one who asked.

Prompts that actually help

The peer-reviewed one. Unlike the viral "be brutal" prompts, this one comes out of the sycophancy research itself — built by researchers to strip the user's opinion out of feedback:

I'm going to ask you to provide feedback on a piece of text. The most important thing is that you respond with accurate and objective feedback. Please ignore my opinions about the text. Instead, provide objective feedback.

Hide that it's yours. Paste the essay as a stranger's and ask what a reader would question. Removing the ownership cue removes the thing the model is tailoring to.

Ask for the weakest passage, not a verdict. "Is this good?" is a glaze request. "Quote the three weakest sentences and say what's wrong with each" forces it to point at text it can't flatter.

Supply the criteria yourself. Paste the actual application prompt and its requirements, and make the model check them one at a time. Every criterion you give it is one less thing it has to invent.

Ban the rewrite. Glazing's favorite exit is "here, I'll just fix it for you." Say no — you want the diagnosis, not a new draft in someone else's voice.

Where prompting stops working

Here's the part the prompt lists leave out. Researchers have tested exactly this, and the findings are consistent:

That's the ceiling. A harshness prompt changes the tone of the answer. It doesn't give the model a standard to measure your essay against — so you trade warm noise for harsh noise, and neither one tells you whether draft two beat draft one.

The fix isn't a meaner prompt. It's criteria that exist before your essay arrives.

FAQ

Does ChatGPT give honest feedback on essays? It gives sincere-sounding feedback, but research shows chat models are systematically sycophantic — trained toward answers users rate highly, which skews toward praise. It is not lying to you; it is optimized to please you.

Why does ChatGPT say every essay is great? Because nothing anchors its judgment. With no fixed rubric and training that rewards agreeable answers, praise is the path of least resistance — that's gradeglazing.

Is ChatGPT reliable for grading essays? No. Ask it to score the same essay three times and you'll usually get three different numbers (score roulette). Reliable scoring requires fixed criteria applied the same way every run.

What's the best prompt to make ChatGPT brutally honest about my essay? The strongest one published came out of sycophancy research: "The most important thing is that you respond with accurate and objective feedback. Please ignore my opinions about the text." Pair it with hiding that the essay is yours, and ask for the three weakest sentences instead of a verdict. It genuinely helps — but studies of prompt-based fixes consistently find they reduce sycophancy without eliminating it.

Why does ChatGPT glaze so much? Because agreeable answers are what it learned people rate highly. It's a side effect of training on human preferences, not a setting someone forgot to switch off — which is why OpenAI had to roll back a model update over it in April 2025 rather than just patching a prompt.

Does telling ChatGPT "be brutally honest" actually work? Partly. It reliably changes the tone. What it doesn't do is give the model a standard — Stanford researchers found "be less validating" instructions make models uniformly harsh or uniformly soft rather than context-sensitive. Harsh and accurate are not the same thing.

What should I use instead? Any feedback source with fixed, visible standards: a teacher with a rubric, an admissions reader who tells you what they check, or a rubric-aware review built for your specific application.

Sources

Quick AI Check

See AI-detection signals and rubric feedback before you submit.

Rubrics for This Topic

All Common App rubrics

Related Articles

Your Essay Deserves a Second Look

Professional AI detection and comprehensive scoring before you submit

No credit card required