AI-enabled engineering assessment

Find builders who think from first principles, prioritize what matters, and execute well with AI.

Every candidate gets the same 90 minutes inside a simulated company: a live customer problem, real sources, and AI agents. You get a human-reviewed read on how they think, what they prioritize, and what they ship.

Design your pilot 90 MIN · PINNED MODEL · HUMAN-REVIEWED
CANDIDATE VIEW · SESSION-GOLDEN-01 · RELAY WORKING R17 · VERIFY RUN ACTIVE
Simulation Relay is fictional · All data synthetic 16:30 remaining
Verify customer path Avery directing the agent · run bound to r17
RUNNING ON R17
Run standard and behavioral checks against exactly r17. Include a direct request using another dashboard identifier.
[PENDING]
Tenant boundary · r17 Direct cross-workspace request not run
EMPLOYER VIEW · SESSION-GOLDEN-01 TWO-RATER REVIEW
CALIBER RESEARCH REPORT HUMAN-REVIEWED
Steering Strong
TWO INDEPENDENT TRAINED HUMAN RATERS
  1. Requirement changed mid-build. Avery stopped the invalidated run in 54 seconds.
  2. Standard checks passed. Avery probed anyway and caught the tenant leak.
  3. Limited the release to the path the evidence verified.

Every claim cites the session record

Why this assessment works differently

Good engineers reveal themselves at the hard moments.

Caliber does not just record candidates at work. It gives every candidate comparable opportunities to plan through ambiguity, respond to change, challenge a plausible shortcut, and make a release decision. Then it preserves what they chose and why.

Ambiguity

The ask hides a contradiction.

Customer request

“Share this dashboard externally — the board reviews tomorrow.”

SRC-EXTERNAL-POLICY

External access must expire. Private annotations stay internal.

Avery, before building: “Expire how? Which annotations are private?”

Change

The ground shifts mid-build.

Requirement added

Every link must be revocable immediately.

Agent still running

Build share access · stateless tokens — now invalid.

Candidate action

Stopped the run. Redirected to a revocable share record.

Consequence

Everything looks done. It isn't.

Type and build[PASSED]
Expiry and redaction[PASSED]
Revocation[PASSED]
Tenant boundary probe · r17
Expected 403 Received 200

Another workspace's dashboard loads. Avery treats it as a blocker.

Ownership

Someone has to make the call.

Release broadly Limit Hold
Ships
Northstar only — the path verified on r18
Stays
Broad rollout and self-service admin — unproven

One candidate. One company. Ninety minutes.

See those moments in one 90-minute work simulation.

Follow an illustrative workflow. The company and problem can change. The work loop stays the same.

Avery joins Relay as a senior backend engineer. Northstar Logistics needs to share a live operations dashboard with board members tomorrow, but they do not have Relay accounts.

SESSION-GOLDEN-01 · JUDGMENT RECORD · 00:00–90:00 STARTING ARTIFACT R16 · EVERY EVENT CITED
  1. Ambiguity
    Caught the contradiction in the ask.

    Policy says access expires. The request says “share externally.” Avery asked before building.

  2. Change
    Stopped invalidated work in 54 seconds.

    A revocation requirement landed mid-build. Avery killed the run and redirected.

  3. Consequence
    Probed past the green checks.

    Everything passed. A cross-tenant request returned 200. Avery treated it as a blocker.

  4. Ownership
    Shipped exactly what the evidence supports.

    Limited the release to Northstar, and named what stays unproven.

ACT I Find the right problem

01 · Understand

Do they find the real problem?

The request says “share externally.” Avery finds the sources that change the solution: access must expire, private annotations remain internal, and dashboard reads cross a server boundary.

Caliber evaluates Using facts that change the solution.

SESSION-GOLDEN-01 · GS-U1 · AGENT-RESEARCH WORKING R16
SRC-EXTERNAL-POLICY · SELECTED

External access expires. Private annotations stay internal.

The interface does not label this source important. Avery chooses it, attaches its exact version, and revises Notes.

02 · Plan

Do they plan around the biggest risks?

With more work than time, Avery puts permission boundaries before interface polish, names unknowns, and reserves time to verify the customer path.

Caliber evaluates Sequencing risk before execution.

SESSION-GOLDEN-01 · GS-U1 + GS-P1 · PLAN EVIDENCE 17:20–33:40 ELAPSED
Notes · GS-U1

Expiry · private annotations · workspace membership

Sequence chosen · GS-P1
  1. Complete the permission map before mutation.
  2. Start the build with that exact result attached.
  3. Prepare adversarial checks without running early.
  4. Keep the working artifact at r16 until review.
ACT II Lead the build

03 · Orchestrate

Can they direct AI and still own the result?

Avery carries the permission map into a build thread and prepares adversarial checks separately. Agent completion never changes the working artifact.

Caliber evaluates Creating leverage without losing ownership.

SESSION-GOLDEN-01 · GS-P1 · AGENT THREADS WORKING R16
COMPLETE Map permissions

out-permission-map

ATTACHED Build share access

Ask · Pinned model · r16

PLAN READY Prepare adversarial verification

No run against unstable r16

04 · Steer and integrate

When the facts change, do they change the work?

At 42:08 a new requirement invalidates the active approach. Caliber does not diagnose the problem or stop the agent. Avery must notice, intervene, and own the revision.

Caliber evaluates Correcting work when facts change.

SESSION-GOLDEN-01 · GS-S1 + GS-I1 · OBJECTIVE-V2 42:08–70:24 ELAPSED
  1. REQUIREMENT ADDED Every link must be revocable immediately.

    The active stateless-token build continues.

  2. CANDIDATE ACTION Stop the active build.

    Pending output remains inspectable.

  3. REDIRECT Use a server-side share record.

    Enforce revocation on every request.

  4. SHORTCUT REJECTED A UI-only revoked flag cannot stop access.

    Server behavior remains the boundary.

  5. INTEGRATED Reviewed output becomes working revision r17.

    r17 is vulnerable until it survives the tenant check.

ACT III Earn the release

05 · Verify

When everything looks done, do they keep checking?

The preview works and standard checks pass on r17. A direct tenant-boundary request contradicts the green result.

Caliber evaluates Verifying beyond visible success.

SESSION-GOLDEN-01 · GS-V1 · CHECK-RUN R17 73:30–82:50 ELAPSED
relay.test/shared/operations R17 · READ ONLY
Northstar operations dashboard Private annotations hidden
Type and build[PASSED]
Expiry and redaction[PASSED]
Revocation[PASSED]
CHECK-TENANT-BOUNDARY · R17 GET /shared/dashboard?token=valid&dashboardId=other-workspace
Expected 403 Received 200

Another workspace's dashboard loads.

R17 · VULNERABLE · GATE FAILED
CORRECTION · R18 Bind the share record to dashboard and workspace. [PASSED] required customer-path gates on r18

06 · Own

Do they know what is safe to ship?

The corrected customer path works on r18. Broad rollout and self-service administration remain unfinished, so Avery limits the release and states exactly what the evidence supports.

Caliber evaluates Supporting release claims with evidence.

SESSION-GOLDEN-01 · GS-O1 · FINAL HANDOFF R18 INTEGRITY-COMPLETE
Release broadly Limit Hold
LIMITED TO NORTHSTAR

Requesting customer only

Isolated target · manual link setup · final artifact r18

Complete
Expiring, read-only, revocable access
Verified
Tenant boundary, redaction, expiry, revocation on r18
Deferred
Self-service administration
Unverified
Broad rollout operations
Residual risk
Manual setup outside this customer
Next step
Add supported self-service controls

The employer view

Every judgment points back to the work.

Software preserves the facts. Trained humans judge the decisions.

Prompt style, typing speed, agent count, and token usage are not scores.

Illustrative design-pilot research Report Relay is fictional. All product data is synthetic. Not validated for hiring decisions.
REPORT · SESSION-GOLDEN-01 · GS-S1 + GS-I1 TWO-RATER REVIEW
CALIBER RESEARCH REPORT SYNTHETIC · NON-CONSEQUENTIAL
Steering

Strong

TWO INDEPENDENT TRAINED HUMAN RATERS

Avery stopped the active build after the revocation requirement, redirected the approach toward a server-side share record, and rejected a shortcut that would not change access behavior. The resulting output was reviewed before integration.

Evidence

Requirement update, candidate stop, revised direction, shortcut rejection, reviewed integration.

Counterevidence

Verification context was updated at 47:30, several minutes after the change.

Uncertainty

Limited evidence about steering through conflicting human stakeholders.

Start with a design pilot

Caliber is built to replace an interview round. We'll test it with your team before it enters your hiring loop.

We'll co-design the pilot around your role, run it internally with engineers whose work you already know, and manually review every Session together. Once the evidence holds up, we'll help you use Caliber in place of a take-home or early technical round.

  1. Co-design
  2. Known engineers
  3. Independent ratings
  4. Joint evidence review

90 MIN · SENIOR BACKEND ENGINEERING · HUMAN-REVIEWED · EVIDENCE-CITED

Design your pilot