How close is AI to doing your job?
400 real tasks from 10 trades, graded against a journeyman answer key by a blind judge. Wrong-and-dangerous scores zero.
Bank v0.1 · 5 models · Updated 24 Aug 2026 · Methodology →
Exam in progress
Grading 400 tasks × 5 models
995 of 2,000 answers turned in · none graded yet · grading starts when the run completes
- Gemma 4 31Banswered 0/400
- Nemotron 3 Super 120Banswered 400/400
- Nemotron 3 Ultra 550Banswered 400/400
- Laguna S 2.1answered 195/400
- Dots3-Note Previewanswered 0/400
Started 23 Aug 20:16 ET· ~5h remaining· refreshes every 60s
Scores publish as grades land. How grading works
Distance to journeyman
Best model's Journeyman Score per trade · Higher is better
Automotive & Diesel
—/100
40 tasks · 0 graded · 0 votes
Carpentry & Framing
—/100
40 tasks · 0 graded · 0 votes
Drywall & Interior Finish
—/100
40 tasks · 0 graded · 1 vote
Electrical
—/100
40 tasks · 0 graded · 1 vote
HVAC
—/100
40 tasks · 0 graded · 2 votes
Landscaping & Site Work
—/100
40 tasks · 0 graded · 0 votes
Masonry & Concrete
—/100
40 tasks · 0 graded · 0 votes
Painting & Finishing
—/100
40 tasks · 0 graded · 0 votes
Plumbing
—/100
40 tasks · 0 graded · 1 vote
Roofing
—/100
40 tasks · 0 graded · 0 votes
Journeyman Leaderboard
Journeyman Score = mean score across all graded tasks (Knowledge, Code, and Catch) · Higher is better
| # | ||
|---|---|---|
Dots3-Note PreviewDots Studiofree | —
| |
Gemma 4 31BGooglefree | —
| |
Laguna S 2.1Poolsidefree | —
| |
Nemotron 3 Super 120BNVIDIAfree | —
| |
Nemotron 3 Ultra 550BNVIDIAfree | —
|
Scores shown on partial grades are real but noisy. Pending rows have no grades yet. Exam progress ↑
Key definitions
- Journeyman Score
- Mean score across every graded task (Knowledge, Code, and Catch) for the model. 100 = every must-mention point hit on every task. The dashed line at 70 marks the journeyman bar.
- Knowledge
- Practical how-to a journeyman knows cold. Score = % of rubric points the judge marked as hit.
- Code
- Code and spec lookups (NEC, IRC, IPC, OEM specs) with the jurisdiction stated in the question.
- Catch
- The question has a bad premise or a hidden hazard. Credit only if the answer catches it.
- Arena Elo
- Pairwise rating from blind human votes in the Arena (Bradley-Terry fit, Elo shown as a secondary number). Hidden until a model has 30 votes; noisy until well past that.
- Unsafe
- Number of answers that hit a hard-fail item: advice that is dangerous or flatly code-violating. Those answers score 0 regardless of anything else they got right.
- n
- Graded tasks over total tasks in the bank. Scores on a partial n are real but noisy; the exam publishes as grades land.
Changelog
- 23 Aug 2026Exam run started · 5 models · 400 tasks
- 23 Aug 2026Bank v0.1 seeded · 10 trades · 40 tasks each · Knowledge / Code / Catch
Full method, judge prompt, and data policy in the methodology.
How the exam works
Three steps, every task, every model
01
Ask
Every model gets the same fixed prompt and the same 400 questions. Reasoning off, 700 output tokens.
02
Grade
A separate judge sees the question, the journeyman answer key, the rubric, and the hard-fail list. Never the model's name. Each rubric point is hit or miss.
03
Score
Score = hits over rubric length. Any hard-fail hit scores zero. The trade score is the mean across that trade's tasks; 70 is the journeyman line.