An open, self-contained portal & benchmark suite to measure any AI model for Telco use - interact with it, observe it, benchmark it. Datasets embedded in the repo within.

Get Started on GitHub Read the Docs
16 telco benchmarks - two tracks LLM-as-judge suites Per-question failure reports Parity-validated scoring Any OpenAI-compatible endpoint MIT licensed
Two-track Benchmark tab - Legacy Test Benchmarks and 2026 Test Suite, two live models

Why TelcoAIBench?

Leaderboard show-ups without a {explicit test suite scopes, verification reports, what platform they have been tested, ...} -> Cannot Be Trusted!

TelcoAIBench packages everything needed to build & reproduce a telco AI model evaluation - the runner, the datasets, and the receipts - inside one repository.

Portal

Expert telco chat personas, a live model identity strip, prompt engineering, and a real-time vLLM observability dashboard. Point it at any OpenAI-compatible endpoint (locally hosted with vLLM and/or AIaaS API) with two environment variables.

Benchmark Suite

The 8 Open-Telco benchmarks plus 2 LLM-as-judge suites, with every dataset embedded as gzipped JSONL. Run from the UI or the CLI - the engine is identical, single-file, stdlib + requests only.

Receipts

Scoring parity- aligned with the official GSMA Telco Eval harness (within 1 point on all 7 leaderboard tasks), plus a full leaderboard-claim verification report showing why reproducibility discipline matters.

The Portal

One Benchmark Portal, seven tabs: Models, Chat, Prompt Manager, Observability, Benchmark, Leaderboard, Tenants.

Expert telco chat

Multi-persona conversations - Telco, Network, Cloud, Storage experts or your own - with persistent shareable sessions, auto-streaming for large contexts, and document upload. The live hero strip shows the connected model's status, latency, and context window at all times.

Chat tab
Observability tab

Live observability

The dashboard polls the model server's metrics endpoint directly: request rates, latency, token throughput, cache utilization, health, and efficiency analysis - with Plotly visualizations and automated collection.

Multi-tenant quotas

Admins create tenant accounts scoped to allowed local models, with lifetime token quotas (prompt + completion) and benchmark attempt quotas for the AUTO-SCORED suites. Each local model carries its own GPU token pool; tenant traffic charges both, and external AIaaS endpoints stay admin-only.

Tenants tab - accounts, quotas and model pools
Per-tenant token usage build-up in Observability

Per-tenant usage build-up

The Observability tab's quota panel tracks every tenant's token build-up across chat and benchmark traffic - next to the live vLLM metrics of every registered endpoint, side by side.

The Benchmarks

Ten suites, one engine. Provision any number of target models - each gets its own card, two run side by side, and every run can be hard-stopped in seconds.

SuiteTasksScoring
Open-Telcoteleqna teletables telemath telelogs 3gpp oranbench srsranbench 6g_benchDeterministic - ported 1:1 from the official harness, parity-validated
Telco's Last Examtelcos_last_exam - 30 questions, 8 domains, 3 difficulty tiers, 246 pointsLLM-as-judge vs. machine-verified answer keys with grading notes; points-weighted, per-domain breakdowns
Vendor GenAIvendor_genai - 24 deep-dives: 6 vendors x 4 domains, honesty traps includedPer-criterion judging (accuracy 0.40, honesty 0.25, completeness 0.20, depth 0.15) with fact anchors and fabrication bait

Judge Models

Judged suites have no deterministic scorer - a judge model you provision grades every answer and goes on record with the results.

Provision once, judge everything

The dedicated judge form takes an endpoint URL, API key, and model name - for example api.openai.com with GPT-5, or any frontier-class endpoint. Judge endpoints are stored separately and never become benchmark targets.

  • Leave the model name empty and the endpoint's models are discovered into a dropdown
  • Structured JSON grading: exam answers graded against the answer key with what-was-missed lists; vendor answers scored per-criterion
  • Judge recorded alongside every score - runs with different judges are never mixed
  • Lab models can judge each other - any provisioned target is selectable too
Judge model provisioning

Leaderboard

Every clean, full-set run recorded in the portal can be published here as a versioned snapshot. Composite = importance-weighted mean (judged suites weigh heaviest); ranking requires 65% weight coverage, full-set runs only, and one consistent judge.

Coverage chips. AUTO-SCORED - benchmarked on the 8 objective suites (TeleQnA, TeleTables, ORANBench, srsRANBench, TeleMath, TeleLogs, 3GPP-TSG, 6G-Bench): multiple-choice / exact-answer questions scored by a parity-validated machine parser - no judge model involved, fully reproducible. FULL + JUDGED - additionally completed the two LLM-as-judge suites (Telco's Last Exam, Vendor GenAI deep-dives), where a judge model grades free-form answers against unpublished answer keys; these weigh heaviest in the composite and act as the contamination check. IN PROGRESS - suites still being recorded; unranked until weight coverage reaches the ranking bar.
Two composite columns, and why. A weighted mean is only meaningful against a stated set of suites. Composite covers all 10 and is therefore shown only for models that have completed both judge-graded suites - everywhere else it reads –, because a 10-suite composite computed from 8 suites is not a lower-confidence version of the same number, it is a different measurement. Auto-8 covers the 8 machine-scored suites and is defined for every model here, so that column is comparable from top to bottom - which is why it is comparable for every model. The bar tracks each row's ranking metric, so bar lengths descend with rank. The Δ beside a composite is that model's composite minus its own Auto-8: how far submitting to the unpublished judged suites moved it.
What judging costs. Ordering is verified first, then the rest by Auto-8. That is a deliberate choice, not a score: an unjudged model is untested against unpublished questions, and on the models tested so far judging has moved the composite by -0.24 to +0.05, with the largest drops on telecom-specialised models. Until a model has faced that, we will not rank it above one that has. Models below the judging gate (TeleQnA < 0.700) stay in the provisional band permanently. A number carrying * with a hatched bar covers only part of its column's suite set and is not on the common basis.

Our Point of View

Telco AI is entering its AI Grid era: inference is leaving the central cloud and distributing across a five-tier fabric that runs from the radio site to the core data center. Every tier has a different latency envelope, hardware budget, and workload class - and the industry question is no longer "which model is best?" but "which model, at which tier, for which task, under which constraints?"

Placement needs measurement

Tier envelopes are illustrative; placement decisions require workload-specific benchmarking. Parameter count alone tells you nothing about whether a model earns its latency budget on 3GPP prose, RAN logs, or link-budget math. TelcoAIBench measures exactly the three axes placement logic needs: accuracy per telco workload class, latency & token economics, and deployment reality (VRAM, precision, serving stack).

Claims need receipts

Self-reported leaderboard numbers routinely diverge from pinned reproduction - our verification report and the TelecomGPT-R1 claims review document exactly how protocol choices move scores by double digits. Every number here ships with its harness version, prompts, precision, and per-question transcripts.

One harness, every model

All models run the identical pipeline: same prompts, same greedy decoding, same serving stack, same datasets embedded in this repo. Judged suites built from unpublished questions act as the contamination discriminator that public benchmarks cannot provide. Comparable numbers, or no numbers.

AI Grid Tier Fitment

Where each benchmarked model belongs on the AI Grid's placement tiers - derived from measured accuracy, decode speed, verbosity, and VRAM footprint, not from parameter count. Each model is shown at the closest-to-the-edge tier it measurably fits - any model can also serve from deeper tiers, so the edge-most viable placement is the decision-relevant one. Tier 0 (RAN-embedded, <10M params on CPU/NPU) excludes generative LLMs by design.

Failure Reports

Every judged run ends with per-domain, per-difficulty, per-vendor and per-criterion breakdowns - and a downloadable report that shows exactly which questions failed, how, and at what level.

Run report with per-question verdicts

Every question, worst first

The run report (HTML + markdown, saved to the persistent state volume) lists each question with its score, difficulty, verdict, the required elements the candidate missed, and the judge's written rationale - the difference between a number and a diagnosis.

Narrated Walkthroughs

Four short screen-capture tours with voice-over - the portal, benchmarking, judge models, and the CLI. Around one minute each.

1 - The Portal

Live model hero strip, expert telco chat, prompt personas, and the observability dashboard.

2 - Benchmarking

Provision targets, pick suites and tiers, run two models side by side, and hard-stop a live run.

3 - Judge Models

Provision a judge with model discovery, run the judged suites, and read structured graded results, breakdowns, and the downloadable failure report.

4 - CLI & Reproducibility

The same engine from the command line: embedded datasets, judged runs, provenance, and receipts.

Quick Start

Two environment variables between you and a running portal. No source edits.

# 1. serve a model anywhere (example: vLLM)
vllm serve <your-model> --port 8080

# 2. point TelcoAIBench at it
pip install 'gradio>=5,<6' && pip install -r requirements-v2.txt
export SME_API_ENDPOINT="https://my-model-route.apps.mylab"
export SME_MODEL_NAME="my-served-model-name"
python sme-web-ui-v2.py                          # :30180

# 3. or benchmark from the CLI (identical engine)
cd benchmarks/open-telco
python3 otel_eval.py --endpoint https://<model-route>/v1 --model <name>

# judged suites: add a judge
python3 otel_eval.py --tasks telcos_last_exam \
  --endpoint https://<candidate>/v1 --model <name> \
  --judge-endpoint https://api.openai.com/v1 --judge-model gpt-5 --judge-key $KEY