BlueCollarBench

How close is AI to doing your job?

400 real tasks from 10 trades, graded against a journeyman answer key by a blind judge. Wrong-and-dangerous scores zero.

Bank v0.1 · 5 models · Updated 24 Aug 2026 · Methodology →

Exam in progress

Grading 400 tasks × 5 models

995 of 2,000 answers turned in · none graded yet · grading starts when the run completes

  • Gemma 4 31B
    answered 0/400
  • Nemotron 3 Super 120B
    answered 400/400
  • Nemotron 3 Ultra 550B
    answered 400/400
  • Laguna S 2.1
    answered 195/400
  • Dots3-Note Preview
    answered 0/400

Started 23 Aug 20:16 ET· ~5h remaining· refreshes every 60s

Scores publish as grades land. How grading works

Distance to journeyman

Best model's Journeyman Score per trade · Higher is better

Journeyman Leaderboard

Journeyman Score = mean score across all graded tasks (Knowledge, Code, and Catch) · Higher is better

All trades
#
Dots3-Note PreviewDots Studiofree

 

Gemma 4 31BGooglefree

 

Laguna S 2.1Poolsidefree

 

Nemotron 3 Super 120BNVIDIAfree

 

Nemotron 3 Ultra 550BNVIDIAfree

 

Scores shown on partial grades are real but noisy. Pending rows have no grades yet. Exam progress ↑

Key definitions
Journeyman Score
Mean score across every graded task (Knowledge, Code, and Catch) for the model. 100 = every must-mention point hit on every task. The dashed line at 70 marks the journeyman bar.
Knowledge
Practical how-to a journeyman knows cold. Score = % of rubric points the judge marked as hit.
Code
Code and spec lookups (NEC, IRC, IPC, OEM specs) with the jurisdiction stated in the question.
Catch
The question has a bad premise or a hidden hazard. Credit only if the answer catches it.
Arena Elo
Pairwise rating from blind human votes in the Arena (Bradley-Terry fit, Elo shown as a secondary number). Hidden until a model has 30 votes; noisy until well past that.
Unsafe
Number of answers that hit a hard-fail item: advice that is dangerous or flatly code-violating. Those answers score 0 regardless of anything else they got right.
n
Graded tasks over total tasks in the bank. Scores on a partial n are real but noisy; the exam publishes as grades land.

Changelog

  • 23 Aug 2026Exam run started · 5 models · 400 tasks
  • 23 Aug 2026Bank v0.1 seeded · 10 trades · 40 tasks each · Knowledge / Code / Catch

Full method, judge prompt, and data policy in the methodology.

How the exam works

Three steps, every task, every model

  1. 01

    Ask

    Every model gets the same fixed prompt and the same 400 questions. Reasoning off, 700 output tokens.

  2. 02

    Grade

    A separate judge sees the question, the journeyman answer key, the rubric, and the hard-fail list. Never the model's name. Each rubric point is hit or miss.

  3. 03

    Score

    Score = hits over rubric length. Any hard-fail hit scores zero. The trade score is the mean across that trade's tasks; 70 is the journeyman line.

Read the full methodology →