Reproducible LLM Evaluation: NHR-SW

Large Language Models are increasingly evaluated using complex pipelines involving different models, datasets, prompts, evaluation frameworks, and hardware configurations. However, small changes in these components can lead to substantially different results, making it difficult to reproduce and fairly compare published evaluations. The workshop will be held on September 22nd 2026, 13:00 - 16:00.

This online hands-on workshop by NHR-SW introduces practical techniques for conducting reproducible LLM evaluations. Participants will work with the EleutherAI LM Evaluation Harness and Weights & Biases (W&B) to reproduce existing results, build custom evaluation tasks, track experiments, and scale evaluations to larger model sets.

The workshop is organized into four practical parts:

Part 1: Reproducing Existing Evaluations

Learn how to reproduce published or existing LLM evaluation results using the EleutherAI LM Evaluation Harness. We will examine evaluation configurations, model settings, tasks, and metrics, and identify the factors that can affect reproducibility.

Part 2: Standardized & Custom Evaluations

Run standardized benchmarks using the LM Evaluation Harness and learn how to create your own evaluation tasks for new datasets or research questions. Participants will work through the process of defining tasks, configuring evaluation settings, and interpreting results.

Part 3: Experiment Tracking, Logging & Reporting

Use Weights & Biases (W&B) to systematically track evaluation runs, configurations, metrics, and results. We will also cover practical approaches for logging experiments and producing clear, reproducible evaluation reports.

Part 4: Scaling & Accelerating LLM Evaluations

Explore practical techniques for speeding up evaluations and conducting large-scale experiments. This includes evaluating multiple models efficiently, using parallel evaluation, working with closed-source LLMs, and scaling evaluations across large model and benchmark collections.

Reference Repositories:

  • EleutherAI LM Evaluation Harness - https://github.com/EleutherAI/lm-evaluation-harness
  • Weights & Biases - https://wandb.ai/

Register Here

by Israel A. Azime