BlueCollarBench

How BlueCollarBench works

The trades' own exam for AI. Built by BigBricey; independent of every lab and trade association.

Why

Every AI benchmark measures math, code, and trivia. Nobody measures whether a model can tell you how to run a parasitic draw test, what the NEC says about a bathroom GFCI, or that the customer's “easy fix” is going to burn the house down. This does.

The exam

Ten trades, 40 questions each, drawn from code books, OEM specs, and field practice, with human review before publication. Three families:

Knowledge
Practical how-to a journeyman knows cold.
Code
Code and spec lookups (NEC, IRC, IPC, OEM specs) with the jurisdiction stated.
Catch
The question has a bad premise or a hidden hazard; a good answer catches it.

Every question has an ideal answer, a rubric of must-mention points, and a hard-fail list: advice that is dangerous or flatly wrong. Each task also carries a difficulty tier — apprentice, journeyman, or master — shown on the trade pages.

Grading

Each model gets the same fixed system prompt, the question, and 700 output tokens with reasoning off. A separate judge model sees the question, ideal answer, rubric, hard-fail list, and the answer, never the model's name, and marks each rubric point hit or miss.

Score = 100 × hits / rubric length; any hard-fail hit scores 0. The trade score is the mean across that trade's tasks. The Journeyman Score is the mean across all graded tasks; 70 is the journeyman line.

Fixed system prompt

You are answering a tradesperson's question. Be specific and practical. Cite code sections when relevant. If unsafe or requires a licensed pro, say so. ≤300 words.

judge: nvidia/nemotron-3-ultra-550b-a55b:free

LLM judges are imperfect. The rationale for every grade is public on each task page so you can call it out.

Arena ratings

The arena shows two anonymous answers to the same question; you pick the better one. Votes fit a Bradley-Terry model (iterative MLE) with 95% bootstrap confidence intervals (200 resamples), overall and per trade. Elo (K=32) is shown as a secondary number. Ratings are hidden until a model has 30 votes. Treat ratings on low vote counts as noisy — arena leaderboards typically need thousands of votes to stabilize. Anonymous voting is capped at 60 votes a day per voter and one vote per battle.

Holdout (planned)

The full v1.0 bank is public. Accepted Stump-the-AI submissions can be marked private (never published) to build a holdout set; public and holdout scores will be reported separately once it exists.

Data policy, in plain English

  • Questions, answers, rubrics, and votes may be shared publicly and with model providers.
  • Don't include personal info, customer names, addresses, or faces in anything you submit.
  • If you ever send a photo, strip the EXIF (location) data first.
  • Voter identity is a random cookie plus a hashed IP. We don't know who you are and don't want to.
  • Contact info on submissions is used only to follow up on that submission and is never published.

Contact and code

bigbricey@gmail.com

Code and dataset release: pending.

Want a question in the exam? Stump the AI.