JUDGe
"Rigorous evaluation is the backbone of trustworthy AI — let's scrutinize the scrutinizers."
A full-day workshop on building reliable, valid, and robust LLM-based evaluators. We bring together NLP researchers, ML systems builders, safety scientists, and industry practitioners around a single foundational question: how do we know whether an LLM evaluator is actually measuring what we intend it to measure?
About the Workshop
Evaluation validity is not a property of a judge in isolation — it is a property of a judge in a system. A well-calibrated evaluator can fail systematically when deployed in a pipeline where its outputs gate safety decisions, feed back into training, or depend on context it was never designed to handle. The field has treated evaluation as a measurement problem — how accurate is the judge? — when the harder question is infrastructural: how does a judge's error profile interact with what is upstream and downstream of it, and what happens downstream when it fails?
JUDGe is the first NeurIPS workshop to take evaluator reliability and validity seriously as a systems problem. Judges embedded in RLHF, DPO, and Constitutional AI pipelines don't just mismeasure — their failure modes compound into model weights and downstream decisions. This workshop is where NLP evaluation researchers, alignment scientists, and production ML practitioners come together around that shared problem.
Construct validity, calibration, human–model alignment, criteria pre-registration, and inter-judge consistency.
Robustness of LLM judges to meaning-preserving paraphrase, length and formatting bias, and semantic content scoring.
Positional and ordering bias in pairwise evaluation, self-preference, and feedback loop risks in judge-guided alignment.
Semantic drift under iterative transformation, adversarial robustness of safety evaluators, and meaning preservation.
Adversarial benchmark design, criteria drift in datasets, domain-specific frameworks, and negative-result datasets.
Prompt engineering, fine-tuning vs. prompting, multi-judge ensembles, tool-augmented judges, and capability gaps.
Production evaluation pipelines, practitioner–researcher gaps, deployment failure case studies, and cost–quality trade-offs.
Reasoning chain validity, multi-turn agentic evaluation, cross-lingual reliability, and ethical dimensions of automated evaluation.
Research Agenda
LLM judge failures don't arrive in isolation — in production pipelines they cascade. Sycophancy biases preference data, which shifts model style, which drifts the judge's implicit criteria across training iterations. Surface sensitivity enables adversarial safety bypasses. Correlated errors across judge families mean ensembles suppress disagreement precisely on the cases most likely to be collectively wrong. The taxonomy below organizes these into seven empirically grounded failure facets, each with an open question at the frontier of current research.
| Facet | Core failure mode | Open question |
|---|---|---|
| 01 Surface vs. semantic sensitivity |
Formatting and length drive scores over meaning; paraphrases receive inconsistent judgments. | Can judges be calibrated to score meaning-equivalent responses identically? |
| 02 Criteria drift |
Criteria shift after seeing real outputs; "evaluate helpfulness" is operationalized inconsistently. | Can criteria be pre-registered and verified for post-hoc consistency? |
| 03 Positional & ordering bias |
Primacy/recency effects skew pairwise rankings. | Does position-swap averaging fully debias long-context evaluation? |
| 04 Sycophancy & self-preference |
Judges favour stylistically familiar outputs regardless of quality. | How do we detect and break the sycophancy–training feedback loop? |
| 05 Reasoning chain validity |
Judge capability bounds evaluation capability; gap worsens as models outpace judges. | What is the minimum judge–model capability gap for valid evaluation? |
| 06 Safety-relevant meaning preservation |
Judges are fooled by adversarial paraphrase preserving harmful content. | What protocols reliably detect safety-relevant semantic drift? |
| 07 Inter-judge consistency |
Cross-judge agreement is low on semantically complex cases. | What inter-judge agreement threshold is acceptable in high-stakes settings? |
Community Deliverable
One concrete output of JUDGe is a structured disclosure template for judge deployment decisions — analogous to a model card, but for the evaluation pipeline. No such standard currently exists. The organizing team is drafting a seed version from collective production experience at Meta, Amazon, and Google, releasing it on GitHub ↗ before the workshop, and refining it collaboratively with attendees during the poster session and panel. The post-workshop version will be published as an open-access community standard.
The template covers four dimensions:
The judge's training data sources, known capability bounds, and distribution assumptions.
What decisions or training signals the judge's scores produce, and what is upstream and downstream.
Which of the seven failure facets apply, with empirical evidence where available.
What validation was performed before deployment, including human agreement rate and methodology.
Timeline
All deadlines are 11:59 PM AoE. All notifications precede the mandatory NeurIPS deadline of September 29, 2026.
Submissions
JUDGe welcomes original research on all dimensions of LLM evaluator reliability and validity. Works in progress, negative results, practitioner case studies, and cross-disciplinary contributions are particularly encouraged — the workshop is designed for work that wouldn't fit neatly into a standard NLP or ML venue track.
All submissions via OpenReview, double-blind, ≥3 reviews per paper. All accepted work is non-archival and posted on the workshop website (authors may opt out). Previously published work at a major ML venue is not eligible.
Program
Full-day program (~9:00–17:30). At least 40% of scheduled time is allocated to contributed talks, posters, and open discussion.
Invited Speakers
Anthropic
Keynote
Microsoft Research NYC
Panelist
Georgia Tech
Panelist
Amazon Core AI
PanelistOrganizing Committee
Northeastern University (Seattle) & University of Washington
Lead Organizer · Corresponding
Meta & Northeastern University
Amazon
Google Cloud AI
Amazon Ads
Diversity & Inclusion
≥50% of confirmed invited speakers identify as women and/or from underrepresented groups. Outreach runs through WAI, Black in AI, LatinX in AI, and Queer in AI networks.
Dedicated oral slots for student and early-career authors (≤3 years post-PhD), each paired with a senior researcher for pre-workshop written feedback.
Complimentary registrations available for junior attendees from underrepresented groups, funded through external sponsorship. Contact the lead organizer for inquiries.
Cross-disciplinary perspectives, practitioner case studies, negative results, and works in progress are all welcome — not just finished research.
Get in Touch
Questions about the workshop, submissions, sponsorship, or travel support? Reach out to the lead organizer.
Workshop logistics, accessibility, sponsorship