A guide per stage with a rubric of anchors, signals to look for and red flags, plus up to four probes generated specifically from the concerns about the candidate you are about to interview.
Each interviewer asks whatever comes to mind, scores on a scale only they understand, and writes three lines at the end. Comparing two candidates becomes comparing two opinions about different things.
Everyone asks from the same guide, with anchors that say what a weak, adequate and superior answer looks like, and the extra questions are the ones missing for that specific candidate.
For each question, the guide describes what characterizes a merely sufficient answer, a partial one and a superior one. The assessor recognizes the pattern instead of inventing a criterion.
The competencies come from the role’s map: mandatory requirements come in already approved; behavioral and contextual ones are suggested and you confirm them.
Tell me about a product decision you changed after talking to a user. What did you believe before, and what changed your mind?
"Your assessment flags discovery without an impact metric. In the case you just described, what changed in the numbers afterwards?"
If their assessment flagged discovery without an impact metric, the probe asks exactly that. If a section of the test was weak, that section becomes a question.
There are up to four per sheet, in a neutral tone. They are not trick questions, they are the confirmation of what was left open.
Below what the stage expects: becomes an interview probe, not a cut.
"Walk me through the most recent A/B test you designed: the hypothesis, the sample size, and what you decided when the result came back ambiguous."
Adjustments are audited. The score is informative: it never eliminates anyone on its own.
Good questions are an asset. The product treats them as one.
Stages on a canvas with parallel groups and transitions by outcome. External interviewers score through a portal of their own.
See inside →TestsEVO AssessA technical test generated from the role, languages with spoken assessment at CEFR level, and behavioral, with a live defense of the work.
See inside →LearningDashboard and calibrationRejection and shortlist reasons, distribution by dimension and the real outcome at 90 days and 12 months feeding back into the analyses.
See inside →It takes a real role to see EVO work, not a canned example.
Start using it