Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

2 Commits
 
 
 
 
 
 
 
 

Repository files navigation

Common App Essay Evaluator

A harsh, evidence-based rubric for college application essays, designed to be pasted into an LLM.

Most application essays are not bad. They are competent and forgettable — they read fine and reveal nothing. That failure is invisible to proofreading, invisible to your English teacher, and fatal in an admissions committee. This tool is built to catch it.


What it does

Paste SKILL.md into Claude, ChatGPT, or Gemini, then paste your essay. You get back:

  • Three pass/fail gates — disqualifiers, an AI-writing detector, and a prose competence floor
  • A score out of 100 across 13 weighted metrics, with a verbatim quote justifying every one
  • The worst sentence in your essay, named and explained
  • What actually works — never padded, never a compliment sandwich
  • A ceiling diagnosis — the highest score this essay can reach with revision alone, and what would have to change about the underlying material to go past it
  • Concrete line-level rewrites for weak passages, preserving your facts and your voice
  • A ranked transformation plan, ordered so that fixes don't undo each other

It will not tell you whether you will get in. Nothing can.


Quickstart

  1. Open SKILL.md, copy the whole file.
  2. Paste it into a new chat with any capable model.
  3. Send this:
Evaluate this essay.

PROMPT: [which Common App prompt, or the exact supplemental question]
SCHOOL: [optional]

ESSAY:
[paste your full essay]
  1. Then ask for line edits or transform depending on what you want next.

Optional, and worth it: paste your activities list too. One of the 13 metrics measures whether your essay reveals anything your application doesn't already say, and without the list that metric is guesswork.


What makes it different

It is calibrated to be strict. Real drafts score 55–75. The rubric explicitly instructs the model to take the lower score on any tie, to refuse to pad the "what works" section, and to never open with "this is a strong start." An evaluator that flatters you is not helping you submit a better essay; it is helping you submit this one.

Identity is weighted at 40%. Non-substitutability, revelation, depth of motivation, and whether the essay is quietly invoicing the reader for credit. There is a floor rule: if this cluster scores under 24/40, the verdict caps at "Danger Zone" no matter how good the prose is. A beautifully written essay that reveals nobody is not a good essay.

It detects manufactured prose and refuses to score it. Twenty signals in two layers. The first twelve are surface patterns — nominalisation load, sentence-length uniformity, symmetry, aphoristic closers. Those are the easy half, and an essay can score zero on all of them and still read as obviously fake.

The second eight are the ones that matter on a polished draft, because they operate on what the writer chose to remember rather than on word choice: composed vulnerability (a writer labelling their own honesty — "that is the flattering version, the real one is..."), thematically convenient detail (a measurement taken afterwards to make an irony land), the parable vignette (anonymous stranger, tidy kindness, stated moral), thesis-illustration architecture (an argument wearing a story's clothes), the flattering flaw (a confessed weakness that survives as praise on a reference letter), and absence of the uncontrolled — the most reliable of all. Real memory carries debris. If every element in an essay serves the theme, it was not remembered; it was designed. That is what readers mean when they say a piece "feels fake" but cannot point to a sentence.

Above threshold it halts and tells you to rewrite it yourself, badly and fast, from memory.

Two things it is careful about: over-coached human writing fires the same signals, so the finding is "this reads as manufactured," never an accusation. And a high reading matters even when a human wrote every word — because admissions officers have always cross-checked essay sophistication against writing scores and recommendations, long before any of this existed. An essay that is better than you is worse for you than an essay that is you.

It fixes your English without taking over your voice. The line-edit protocol supplies concrete replacements for weak sentences under four rules: preserve every fact exactly, preserve the register, preserve the voice, and cut before you rewrite. Roughly a third of weak lines want deletion rather than repair.

It will never invent a detail for you. Where a rewrite needs a fact you haven't supplied, it writes [ADD: ...] and tells you what's missing. A fabricated specific collapses in an interview and is exactly the failure the rubric exists to detect.

Three hard caps override the arithmetic. If Identity scores under 24/40, the verdict caps at Danger Zone no matter how good the prose is. If any load-bearing element is invented or unverified, the total caps at 60 and the score is labelled UNVERIFIED — a rubric applied to fiction measures nothing, and this holds especially when the invented material is good. And a model cannot score an essay it wrote or substantially rewrote; it has to say so and refuse the number, because you cannot see the seams in something you sewed.

When a human says "this feels fake" and the rubric says 90, the human is right. That rule is written into the skill. Surface metrics are easy to satisfy and the deep signals are easy to miss, so a manufactured essay can pass every checkbox while failing the only test that matters. The instruction is to re-audit, not to defend the number.


The 13 metrics

Cluster Metric Points
Identity (40) Non-Substitutability — the Thumb Test 10
Revelation 9
Why-Depth 11
Intrinsic Motivation Index 5
Relatability 5
Craft (30) Opening 6
Ending 7
Show vs Tell 7
Specificity and Anchoring 10
Structure (18) Focus and Slice Discipline 7
Arc and Growth 7
Prompt Fit 4
Trust (12) Voice Authenticity and Register 7
Application Coherence 5

Bands

Score Verdict
92–100 Exceptional — an officer would argue for this in committee. Rare.
85–91 Excellent — competitive anywhere. Submit it.
75–84 Strong, with an identifiable ceiling.
60–74 The Danger Zone — reads fine, reveals little. The most common band and the most dangerous.
45–59 Weak. Usually needs a new angle, not edits.
Under 45 Failing. Usually a topic problem.

Repository contents

SKILL.md                              The evaluator. Self-contained. This is the product.
reference/essay-craft-principles.md   The knowledge base — what works, why, and who says so.
reference/applying-sideways.md        The governing philosophy, summarised.

SKILL.md needs nothing else to run. The reference/ files are for reading rather than for pasting — useful if you want to understand the reasoning behind a score, or if you are writing rather than evaluating.


Where this comes from

The rubric is not invented. It is a formalisation of what admissions officers and published essay collections already say, turned into something a model can apply consistently.

  • Gen & Kelly Tanabe, 50 Successful Ivy League Application Essays — the 26-mistake taxonomy, the why-ladder, the Thumb Test, and 53 per-essay craft analyses. Includes interviews with a former Dartmouth admissions officer and a former Yale admissions officer.
  • The Harvard Independent, 100 Successful College Application Essays — the "hooks" model, the authenticity-mismatch problem, and the storytelling argument. Includes an essay by Fred A. Hargadon, former Dean of Admission at Princeton and Stanford, who has read roughly 200,000 applications.
  • Applying Sideways, MIT Admissions — the argument that doing things for admissions is self-defeating.
  • MIT Admissions"if you're thinking too much — spending a lot of time stressing or strategizing about what makes you 'look best,' as opposed to the answers that are honest and easy — you're doing it wrong."
  • Published work on machine-text detection, for Gate 2 and Appendix A.

No copyrighted text is redistributed here. Books are cited so you can buy them. Quotations are brief and attributed, for commentary and criticism.


Honest limitations

  • It cannot predict admissions outcomes, and it is instructed to refuse when asked. It does not know your file, your competition, or a given school's institutional priorities in a given year.
  • It is not a detector. Gate 2 reports structural signals. Commercial classifiers work at a statistical depth no static rubric reaches, and they false-positive heavily on dense human prose. A low reading here is not a guarantee of anything.
  • Scores will vary between models and between runs. Treat the findings as the output and the number as a summary of them. If two runs disagree by fifteen points, read both sets of evidence and judge for yourself.
  • A rubric cannot supply material. If your essay's ceiling is 78 because the topic is wrong, the tool will tell you — but only you can pick a better one.
  • Two revision passes, then stop. Over-edited essays acquire a sanded flatness that admissions officers name independently. More editing is not more better.

Contributing

Issues and pull requests welcome, particularly:

  • Additional AI-writing signals with real examples
  • Rubric calibration data — if you scored an essay and the result was clearly wrong, say so and show the essay if you're willing
  • Prompt-specific guidance for supplements at particular schools

Please do not open PRs adding copyrighted material.


License

MIT. Use it, fork it, sell nothing pretending it came from an admissions office.

About

A harsh, evidence-based rubric for college application essays, designed to be pasted into an LLM. Three gates, 13 weighted metrics, AI-writing detection, and concrete line edits.

Topics

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors