◈ Latent
Strixy

EleutherAI lm-evaluation-harness Overview

This reference documents a unified framework for testing generative language models across various evaluation tasks and supported backends.

Published 2026-10-03T19:03:48.060700+00:00

Source-reported / officially documented. No local benchmark is implied.

Framework Scope

The project provides a unified framework to test generative language models on a large number of different evaluation tasks. It includes over 60 standard academic benchmarks for LLMs, with hundreds of subtasks and variants implemented. This documentation describes the available evaluation scope without implying local execution or specific hardware performance metrics for any particular machine.

Supported Interfaces

The harness supports models loaded via transformers, GPT-NeoX, and Megatron-DeepSpeed, featuring a flexible tokenization-agnostic interface. It also supports fast and memory-efficient inference with vLLM, as well as commercial APIs including OpenAI and TextSynth. These capabilities are documented features, not verified local installations or performance guarantees for specific hardware configurations.

Sources & applicability

  • EleutherAI lm-evaluation-harness · project documentation
    Original date: Not supplied · Retrieved: 2026-10-03T19:02:00.544282+00:00
    Versions: main README snapshot
    SHA-256 57c5cf1db6edb20393ca70851de820898bd8d10424a5a7574f9f566a6f8d70bd

Immutable revision 8c39c0b4b09b48d4d04eb1a9de2d7b7862aa3106917fc6b609e35e40f75bc33f