The TSA Hiring Process, Scored: A Structured-Interview Rubric
A behavioral 1-to-5 rubric for structured hiring interviews, with anchors, an annotated example, and a calibration exercise, using the TSA process as the worked example.

The TSA Hiring Process, Scored: A Structured-Interview Rubric
Two interviewers sit through the same structured interview. Same candidate, same five scenario questions, same hour. They walk out a full point apart on the score. One of them is sure the candidate was strong; the other is sure he was vague. Neither is lying, and that is the problem. The disagreement is almost never about what the candidate said. It is about what each person remembered forty minutes later.
Most people who search the TSA hiring process are candidates trying to pass it, and there are plenty of guides for that. This one is for the other side of the table: the hiring manager who has to turn an hour of talking into a number two people can defend. The TSA is the sharpest public example of that job. Its airport assessment includes a structured interview where every applicant faces the same five scenario-based questions and is scored on demonstrated competencies, not on how likable they were. If you are hiring anyone, that is the model to steal, and this piece hands you the rubric.
A structured interview does not make interviewers smarter about people. It makes them consistent. The federal government has run on this for decades: the Office of Personnel Management's structured-interview method asks every candidate the same questions in the same order, and scores every answer against the same rating scale and the same standard for a good response. The research behind it is not subtle. Interview structure raises predictive validity along a straight line, from about .20 with no structure to about .57 at the top (Huffcutt & Arthur, 1994); a 2022 re-analysis puts structured interviews at .42 against .19 for unstructured ones (Sackett et al., 2022). Structure roughly doubles what an interview can predict, and the mechanism is boring: same questions, same rubric, scored from the words.
Scored from the words is the part that breaks in practice. A rubric read from memory after lunch is a vibe with a number stapled to it. This piece scores an interview the way the TSA scores one: six criteria, each with what a weak answer sounds like set against what a strong one sounds like, so two people land on the same number. It is also how talk2bud reads an interview. The app captures the call from your Mac's system audio, with no bot joining the meeting, and reads the transcript against an Interview Lens that returns the same six scores with the exact line the candidate said to earn each one. No bot in the room is a capture choice, not a license to record in secret: you still tell the candidate you are recording, which is both the honest version and, for an interview, the fairer one. The rubric below and that Lens are the same object.
Everything here was verified in August 2026 against the sources linked inline. One honest limit up front: the TSA keeps its internal scorecard private, and this piece does not pretend to have it. The rubric is a general structured-hiring rubric, with the TSA process as the worked example.
The One-Page Play
The method fits on an index card. Take your last interview. Score each of the six criteria a 1 to 5 based only on what the candidate said, and write the line that earned the score. The two lowest criteria are what your next round should probe. A criterion you cannot attach a quote to is a criterion you did not actually score.
Watch one number above the average: consistency across questions. Score all five questions separately, then look at the spread. A candidate who tells one vivid story and four thin ones is not a strong candidate who had an off moment. He is a candidate with one rehearsed story, and averaging hides it. The full playbook, with anchors for every score band, a calibration exercise, and a 30-day rollout, is in the download section below, as a PDF and an importable Lens.
The Structured-Interview Rubric
A checklist asks "did the candidate give an example?" and any answer with the word "example" in it ticks the box. A rubric asks what a 2 sounds like against what a 5 sounds like, with the phrasing attached, so the score survives a second interviewer. That is the property the federal structured-interview method is built on, and it is the difference between a scorecard and a mood.

Here is the full rubric, in a form you can copy into your own notes. Score each row 1 to 5, and quote the line that earned it.
| Criterion | Listen for | A 1-2 sounds like | A 4-5 sounds like |
|---|---|---|---|
| 1. A real situation, not a maxim | A specific past event, with a when, a where, and who was there | "I always stay calm and professional under pressure." | "Two weeks ago on the late shift, a customer demanded a refund we stopped giving after 30 days." |
| 2. Their action, owned as "I" | What this person did, in the first person, not what the team did | "We handled it as a group and got through it." | "I checked the receipt date myself and gave him the answer before my manager stepped in." |
| 3. A decision with a tradeoff named | A choice between options, and why they picked one | "I just followed the procedure I was trained on." | "Two choices: bend the rule or offer store credit. I offered credit, since a chargeback costs us more." |
| 4. A result you can check | An outcome with a number or an observable consequence | "It worked out fine in the end." | "He took the credit, and we made that the standard reply for late refunds afterward." |
| 5. What they would do differently | Honest reflection, not a humblebrag | "Honestly, I wouldn't change a thing." | "I'd have brought my manager in earlier so he heard the policy from both of us." |
| 6. On the competency the question asked | Evidence for the thing you asked about, not a nearby virtue | Asked about conflict, answers about being organized. | Asked about conflict, the whole story is about the conflict and how it resolved. |
Two of these are hardest to fake and easiest to verify: a real situation (criterion 1) and a checkable result (criterion 4). A candidate can rehearse ownership language and a tidy line about what he learned. What he cannot manufacture on the spot is a dated, located situation and a result you could go check. When you are deciding whether you scored real evidence or scored a performance, those two are what tell you.
How a Good Answer Actually Sounds
The rubric only means something against real phrasing, so here are two answers to the same scenario question. They are illustrative examples written to show the scoring, not a real transcript, and dated the way a real review would be: an interview from 12 August 2026. The margin notes are what the rubric writes.
The question, a standard past-behavior prompt: "Tell me about a time you had to enforce a rule that upset someone."
First, a weak answer.
Them: I think staying professional is really important. If someone's upset, I stay calm, I explain the policy, and I make sure they feel heard. I'm a big believer in customer service, so I'd de-escalate and try to find a solution that works for everyone.
Criterion 1 scores a 2: there is no situation here. "If someone's upset" and "I'd de-escalate" are how the candidate imagines himself behaving, not a thing that happened. With nothing resolved because nothing occurred, criterion 4 scores a 1, since there is no result to check. The answer is fluent and completely unscoreable on evidence, and the follow-up that saves it is one sentence: "Can you walk me through a specific time that actually happened?"
Now a strong answer to the same question.
Them: Last winter at the checkpoint, a passenger refused to take off his belt and the line was backing up behind him. I had two options, argue it out at the belt or move him, and I moved him to a private screening because the backup was the bigger risk to the whole lane. I explained the rule once, offered the private option, and he took it. The lane cleared in about two minutes, and my supervisor started sending belt refusals to private screening as the default after that. If I did it again, I'd have flagged my lead a minute sooner instead of trying to hold both jobs myself.
This one scores a 5 almost across the board, and you can say why. Criterion 1 is a dated, located, specific event. Every verb is "I," so criterion 2 is a 5 as well. The named tradeoff and the reasoning behind the choice earn criterion 3, the checkable result and its lasting consequence earn criterion 4, and the honest reflection at the end, not a rehearsed weakness, earns criterion 5. Same question, same length of answer, and a completely different candidate underneath, which the rubric lets you prove by pointing at the words.
Calibration: Two Interviewers, One Number
A rubric earns its keep only if two people scoring the same interview land within a point of each other. When they do not, the argument is almost always about an adjective, not a fact. Calibration fixes that by forcing the argument down to the quote.
Try it. Score criterion 4, a checkable result, on this answer, then read the key.
Them: So I took the lead on it, and honestly it went really well. The customer was happy by the end, and my manager said it was handled great. I'm proud of how that one turned out.
Score it before reading on. The answer is a 2. An interviewer who liked the candidate wants to give it a 4, because it sounds confident and there is a manager's praise in it. But criterion 4 is not "did the candidate feel good about the outcome." It is "is there a result you can check." "It went really well" is a feeling, "my manager said it was great" is secondhand, and neither is a number or an observable consequence. What earns the score is the evidence in the words, not the confidence in the voice. Scoring the words instead of the delivery is the whole reason two interviewers agree, and it is the same reason a Lens and a person reach the same score: both are reading the transcript, not remembering the room.
The rule that keeps calibration honest is the quote rule. No score without a line from the interview attached. A score with no line under it is a number you made up, and the moment the disagreement becomes an argument about the quote instead of the impression, it resolves itself. This is the same discipline that makes evidence beat adjectives in a written employee self-assessment: a checkable fact wins, a personality word does not.
How the Rubric Shifts by Role and Question Type
The six criteria hold across hiring, but the weights move.
By question type, the rubric bends. A past-behavior question ("tell me about a time") is scored exactly as above: criterion 1 is the whole foundation, because if there is no real event, there is nothing to score. A situational or hypothetical question ("what would you do if a passenger became aggressive") has no past event by design, so criterion 1 stops applying and criterion 3, the decision and its tradeoff, carries the weight. The OPM method uses both kinds on purpose; the mistake is scoring a hypothetical for a specificity it was never going to have.
By seniority, the burden changes. For a frontline, high-volume role, the kind the TSA hires by the thousand, criterion 6, consistency on the competency asked, matters most, because you are hiring for a floor of reliable judgment across five questions, not for one brilliant story. For a senior or leadership hire, criterion 3 dominates, because naming a tradeoff and owning the call is the judgment that seniority is supposed to buy. A senior candidate who scores high on polish and low on criterion 3 is telling you he has never actually had to decide anything hard.
The Four Ways Interviewers Score This Wrong
A rubric removes most bias, and four failure modes survive if you let them. Structured interviews are meaningfully less prone to these than a free-flowing chat, but "less" is not "immune."
Recency. You remember the last strong answer and let it color the whole interview. Fix it by scoring each of the five questions in order, before forming an overall read.
Halo. One vivid story lifts all six criteria. A candidate can be a 5 on the question he prepared for and a 2 on the other four, and the average will lie about him. Score each question on its own evidence, then read the spread rather than the mean.
Leniency and severity. Two interviewers who run hot and cold produce scores that cannot be compared. This is exactly what the anchor text is for: when a score has to point at the sentence a 2 or a 5 sounds like, personal strictness stops mattering, which is the entire reason behaviorally anchored scales beat "excellent, good, fair."
Similarity. You score the candidate you would grab a coffee with. A panel and a fixed rubric are the defense, because both make you say which competency, in which answer, earned the number, rather than "good culture fit."
Rolling It Out in 30 Days
Week one is a pilot. Score your last five interviews, or your next five if you did not record the last ones, without changing how you run them. You want a baseline and your panel's weakest criterion before anyone gets defensive about a number.
Week two is calibration. Have two interviewers score the same interview independently, then compare. Anywhere you disagree by more than a point, argue it to the quote until the anchor settles it. If you cannot both point at a line, the disagreement is about the person, not the answer, and that is the exact bias the rubric exists to catch.
By week three, every panel scores every candidate on the rubric, and you track one thing: the gap between the average score and the actual hire decision. When you hire the 3.1 over the 4.2, write down why in a sentence, because that sentence is either your real criterion or your bias, and you want to know which.
The final week is for reading the trend, not the averages. If your whole panel scores soft on criterion 3, your interviews are rewarding likable people who never demonstrate judgment, and the fix is a better question, not a sterner scorer. A rubric your team has argued with and edited is a rubric your team will actually use.
Running the Rubric on the Interview You Just Had
You will not score interviews by hand forever, and a panel that runs six candidates a week cannot. Every read is a day old by Friday, and Friday is when the debrief happens. Here is the path from a live interview to a score.
1. Capture the interview without a bot in the room. talk2bud records from your Mac's system audio, so nothing joins the call as a participant and no bot appears in the attendee list. That is a capture choice, not a secrecy one. No hidden bot is not the same as recording in secret, so you still say it out loud to the candidate, which for an interview is a fairness gain as much as a legal one: a recorded, rubric-scored interview is one you can defend to the candidate and to a regulator. The honest line is short: "I record these so I score everyone on the same rubric rather than on memory, is that alright with you?" The system-audio route captures the remote call, whether that is Zoom, Meet, or Teams; how that capture works, and why it beats a screen recorder that loses the other side's voice, is worth a read on its own.
2. Pick the Interview Lens, and make it this rubric. In Modes, switch to Lenses and choose the built-in Interview lens, then duplicate it and import the rubric from this page as its playbook. The Lens is the rubric: the same six criteria, the same quote rule, the same anchors.

3. Read the score, not the recording. When the interview ends, talk2bud scores the transcript against the Lens and hands back the six criteria with the quote that earned each, plus the red flags and the areas the next round should probe. You read a page, not a fifty-minute recording, and the debrief starts from the same evidence for every interviewer instead of five different memories. The candidate who scored lowest on the competency you care about is the one the next round goes after.
That last part is the reason to run it on the machine instead of in your head. A person forgets whether the strong answer came on question two or question five by the time the interview is over. A Lens reading the transcript does not, and it quotes the line back to you before the debrief turns into an argument about who remembers what.
Download the Structured-Interview Playbook
The full playbook is free, under a CC BY 4.0 licence, so you can adapt it and re-share it with attribution.
- The Structured-Interview Playbook (PDF). The thesis, all six criteria with anchors for every score band, the annotated example, the calibration exercise, and the 30-day rollout. Download the PDF.
- The importable Lens (Markdown). The same six criteria as a file talk2bud imports as a call Lens. In the app, go to Modes, then Lenses, then New, and import it. Download the Lens file.
There is no email gate. The playbook and the Lens are the same criteria, so the score you give an interview by hand and the score the app gives it come from the same rubric.
Frequently Asked Questions
How long does the TSA hiring process take?
The TSA describes its Transportation Security Officer hiring process as taking about 90 days on average, though the actual time varies with airport staffing, appointment availability, medical documentation, and background-check complexity. The steps run from a USAJOBS application through a computer-based test, a conditional offer, and an airport assessment that includes the structured interview, a color-vision test, a medical evaluation, and a background investigation. Because timelines depend on the specific airport, many applicants report the full process running three to six months rather than exactly 90 days.
What is a structured interview, and why does it score better?
A structured interview gives every candidate an identical set of predetermined questions and grades each answer on a fixed scale with defined standards, instead of letting each interviewer improvise. It predicts job performance better because that consistency strips out the memory, rapport, and bias that otherwise drive the number. Across meta-analyses, validity climbs as interviews get more structured, reaching roughly .57 at the highest levels (Huffcutt & Arthur, 1994); a 2022 re-analysis estimates structured interviews near .42 versus about .19 for unstructured ones (Sackett et al., 2022).
How do you score a structured interview?
You score a structured interview by rating each candidate on a small set of job competencies, usually four to six, using a fixed scale (commonly 1 to 5) where each score band has a written behavioral anchor describing what that level sounds like. The federal method records a separate score per competency and attaches evidence to each rating, so a score is defensible by the exact response that earned it. The rule that makes it reliable is that no score is given without a specific line from the interview to justify it, which is what lets two interviewers reach the same number.
How many questions are in the TSA interview?
The TSA's standardized Transportation Security Officer interview is commonly described as five scenario-based questions, conducted by two interviewers in a panel format, running about an hour. The questions ask candidates to describe past situations, the actions they took, and the outcomes, and each answer is scored against the competencies the role requires. Prep guides that track the process report candidates generally need to score at least a 3 on each competency to pass, though the TSA keeps its exact scoring criteria private.
What makes a good interview answer?
A good interview answer describes a specific past situation, the action the candidate personally took (in the first person, not "we"), a decision they made with the tradeoff named, and a result you can actually check, such as a number or an observable consequence. Answers built on maxims ("I always stay calm") or hypotheticals ("I would handle it professionally") score low because there is no evidence to verify. The strongest answers also add what the candidate would do differently, which shows reflection rather than a rehearsed script.
Verified August 2026 against TSA's public careers pages, OPM's structured-interview guidance, and interview-validity meta-analyses (Huffcutt & Arthur, 1994; Sackett et al., 2022). The TSA does not publish its internal scoring criteria; the rubric here is a general structured-hiring tool, illustrated with the TSA process. The interview exchanges are illustrative, written to demonstrate the rubric rather than transcribed from real interviews.