How to Write a Rubric: A Step-by-Step Guide for Teachers
Last updated:
A rubric is the single most useful grading tool a teacher can build, and also one of the easiest to build badly — a grid of vague adjectives that took an hour to make and doesn’t actually make scoring faster, fairer, or clearer to students. This guide walks through building one properly: what a rubric is and why it’s worth the up-front time, the basic structural choice between analytic and holistic formats, and a six-step process — from defining the learning goal through piloting the finished draft — for writing one that holds up across a real stack of student work.
What a rubric is and why use one
A rubric is a scoring guide that spells out, in advance, what you’re grading (the criteria) and what each level of quality actually looks like (the descriptors). Instead of carrying a private mental checklist that can drift over the course of marking thirty papers, you write the checklist down once and score every student against the same descriptions.
Susan Brookhart’s 2013 book for ASCD, “How to Create and Use Rubrics for Formative Assessment and Grading,” is one of the standard references teachers turn to for rubric design, and it frames a good rubric around two parts working together: criteria that name what matters in the work, and descriptions of performance at each level along those criteria. Leave out either part — criteria with no real descriptors, or descriptors that don’t map to a clear criterion — and you end up with something that looks like a rubric on the page without functioning like one when you actually grade.
The payoff for that up-front work is threefold. Scoring gets faster once the rubric exists, because you’re matching evidence in the work to a descriptor instead of composing a fresh judgment for every paper. Scoring gets more consistent, both across a large stack and across two teachers grading the same assignment. And expectations become visible before students start writing, not after they get a grade back, which is what turns a rubric into a teaching tool and not only a grading one.
Analytic vs. holistic rubrics
Before you draft rows and columns, decide on a structure. An analytic rubric scores each criterion separately — thesis, evidence, organization, mechanics — with its own descriptors, then adds the rows into a total. A holistic rubric skips separate rows entirely and matches the whole piece of work to one overall description per level.
Analytic rubrics take longer to build and to score, but they hand back richer, criterion-specific feedback, so a student knows exactly which row to work on next time. Holistic rubrics are faster for a single overall judgment — useful for a quick formative check or a low-stakes assignment — but they tell a student less about what to fix. Our guide to analytic, holistic, and single-point rubrics covers the fuller comparison. Everything below assumes an analytic structure, since it’s the default most teachers reach for.
Step 1: Define the task and learning goals
Start from the learning goal, not from a blank grid. Write one or two sentences describing what you actually want students to demonstrate — not “write an essay” but “construct an argument supported by textual evidence and organize it into a coherent multi-paragraph structure.” Naming the target before you build the measuring tool is the same backward-design move behind planning a whole unit: decide what evidence of learning would look like, then design the assessment around it.
From that goal, ask what would separate a response that clearly met it from one that didn’t. That answer becomes the raw material for your criteria and, later, your descriptors — which is why skipping this step tends to produce rubrics that measure whatever the assignment happened to contain, rather than what you actually taught.
Step 2: Choose criteria
Criteria are the rubric’s rows — the distinct qualities you’re going to score. Good criteria are drawn directly from the learning goal, stay as independent of each other as you can manage (so a weak thesis doesn’t automatically sink the mechanics score), and stay few enough that you can hold all of them in mind while grading. Four to six criteria is a practical ceiling for most classroom assignments; an essay rubric, for example, typically scores something like:
- Thesis or claim
- Use of evidence
- Organization
- Style or voice
- Grammar and mechanics
Resist folding in things you’re not actually trying to measure — neatness, length, or effort — unless they’re genuinely part of the learning goal. Brookhart is specific on this point: blending achievement (what the student demonstrated) with behavior or effort in a single score muddies what the grade is actually certifying, and it’s one of the fastest ways to make a rubric feel unfair.
Step 3: Choose performance levels
Levels are the columns — how many degrees of quality you’ll distinguish, typically running from “not yet meeting expectations” up to “exceeds expectations.” Four levels is the most common choice in K–12 practice, largely because it removes a safe middle option and forces a real judgment call on every criterion, while three levels can be enough for a quick formative check and five or six suit high-stakes assessments where finer distinctions are worth the extra scoring time.
How many levels actually serve your purpose, and how to label them so the label communicates a target rather than just a number, is covered in more depth in our guide to how many performance levels a rubric should have.
Step 4: Write descriptors that are parallel, observable, and distinct
Descriptors are the text inside each cell — the part that separates a rubric from a checklist with points attached. Three properties make descriptors work:
- Parallel — describe the same underlying feature across every level in a row, varying only the degree of quality, instead of switching to a different feature at the top level.
- Observable — describe what a reader can point to in the work itself, rather than an inference about the student’s mindset or effort that nobody but you can see.
- Distinct — make each level different enough from its neighbor that two people scoring the same paper land on the same one; vague qualifiers like “somewhat” or “mostly,” with no concrete anchor attached, let adjacent levels blur together.
Heidi Andrade’s research on rubrics and student self-assessment — built up over roughly two decades — keeps landing on the same requirement: descriptors only help students revise their own work when the wording is concrete enough for a student to apply it without a teacher translating it for them. That’s a useful test to run on your own draft: if you can’t picture a student using a descriptor to check their own work, a colleague grading with it will struggle just as much.
Step 5: Weighting
Decide whether every criterion counts equally or whether some matter more for this particular assignment — weighting “use of evidence” above “mechanics” on an argumentative essay, for instance. The simplest approach, and the right default unless you have a specific reason otherwise, is equal weighting. The next simplest is giving each row a different point total (evidence out of 30, mechanics out of 10) rather than applying a percentage multiplier after the fact.
Whatever scheme you choose, keep it something you could explain to a student or a parent in one sentence. A rubric whose final score needs a spreadsheet to reconstruct has quietly given up the transparency that was the point of building one in the first place.
Step 6: Pilot and revise
Before handing a new rubric to a full class, test it against two or three real or realistic samples — ideally one strong response, one average one, and one weak one.
- Score each sample cold, noting which level you assigned per criterion and why.
- Check whether the descriptors actually separate the samples — if the strong and average paper land in the same cell on a given row, that row’s wording needs a sharper anchor.
- Revise the wording that caused hesitation, then reread the row top-to-bottom to confirm it’s still parallel.
- If you co-grade with a colleague, have them score the same samples independently and compare; gaps in agreement point directly at the descriptors that need work.
A first-draft rubric is rarely the version you keep, and that’s normal rather than a sign you did it wrong. Brookhart’s own guidance on rubric quality makes the same point: the biggest gains in scoring reliability tend to come from revising a rubric after trying it on real student work, not from getting the wording perfect before anyone has used it.
Common mistakes
- Vague, subjective language — words like “creative” or “good effort” with no visible anchor in the work itself, so two raters (or the same rater on two different days) can disagree.
- Uneven levels — a crisp top and bottom level with mushy, near-identical wording in between, so the middle of the scale doesn’t actually discriminate between papers.
- Too many criteria — a twelve-row rubric looks thorough on paper and is unworkable in practice; almost nobody holds twelve independent judgments consistently across thirty papers.
- Grading behavior instead of learning — folding participation, neatness, or on-time submission into an academic criterion, which blurs what the score is actually certifying.
- Skipping the pilot — using a brand-new rubric cold on the whole class instead of testing it on a few samples first.
- Rebuilding from scratch every time — reusing a stable rubric across an assignment series, adjusting only the prompt-specific detail, is what makes rubrics a time-saver; starting over each time gives that back.
Build one now
You don’t have to build any of this from a blank grid. Start from a ready-to-edit blank rubric template and swap in your own criteria and descriptors, or start from the essay rubric if you’re grading writing specifically — both already follow the parallel, observable, distinct structure described above, so you’re editing wording rather than inventing structure from scratch.