Why it is written this way
When teams rush to build an interview scorecard, they often end up with a list of generic traits next to a 1–5 scale. Such forms are easy to fill out, but when one interviewer gives a 4 and another gives a 2 for the exact same answer, no one can explain why. The score ends up being nothing more than gut instinct disguised as numbers.
The core strength of this prompt lies in the Task instruction requiring descriptive behavioral benchmarks for 5, 3, and 1 points rather than bare numbers. Instead of relying on their personal intuition, interviewers match candidate responses against predefined descriptions. This drastically reduces scoring discrepancies and makes differences easy to diagnose.
The Role is intentionally set as a hiring manager aggregating multi-interviewer scores. An individual interviewer rarely feels the pain of inconsistent scorecards; only someone responsible for reconciling conflicting ratings understands why rigorous calibration matters.
Making the criterion count a variable is also deliberate. Left unchecked, AI models tend to produce 8–9 criteria. For a scorecard that needs to be completed within 3 minutes post-interview, too many fields lead to rushed, careless scoring at the end. Limiting it to 4–6 allows interviewers to evaluate each item thoroughly.
In the Format, percentage weights ensure that critical core competencies aren't diluted by secondary traits. The three-line post-interview debrief—especially "items unverified"—sets a clear agenda for subsequent interview rounds or reference checks. Finally, the Constraint forbidding vague buzzwords prevents the model from falling back on meaningless placeholders like "Great / Good / Poor".
Unfamiliar terms? See Aha AI: output-format, role-prompting
Compared with a bad example
Make an interview evaluation scorecard.
This generates a generic table with 5 vague criteria and a 1–5 scale. While it looks usable on the surface, the numbers lack concrete definitions, leaving interviewers to use their own subjective yardsticks. If a candidate is rejected later, numbers are the only record left without any actionable evidence. It also tends to include irrelevant generic traits like "Good personality" or "Growth mindset," which turn the scorecard into a channel for pure bias.
Variations
When Interviewer Scores Diverge
The interviewers' scores for a {{hiring job role}} candidate have diverged significantly across multiple criteria. Instead of simply averaging the scores, design a structured 30-minute calibration debrief protocol for the interview panel to reach consensus. Outline the step-by-step agenda, specific questions to ask each interviewer to surface their underlying evidence, and the tie-breaking criteria if consensus cannot be reached.
Averaging split scores often results in advancing lukewarm candidates no one is truly confident in. This variant shifts the focus from blending numbers to surfacing concrete evidence.
For Take-Home Assignments and Portfolio Reviews
Create an evaluation rubric for reviewing take-home assignments or portfolio submissions for a {{hiring job role}} position. Include {{evaluation item count}} evaluation criteria, defining each by elements directly observable in the submitted work. Describe 5-point, 3-point, and 1-point benchmarks through concrete artifact characteristics, and exclude unverifiable factors such as estimated time spent.
Interviews evaluate verbal responses, whereas assignments evaluate tangible deliverables. Because the medium differs, the scoring benchmarks must be anchored directly to work artifacts.
Related prompts
Last updated 2026-09-02 · Found a mistake? Let us know