Haider Ali · Founder & CEO of ClearZone · AI & CS at King’s College London · London

Projects · 02

qecgen — datasets for quantum decoders

By Haider Ali ·

qecgen, by Haider Ali, is an open-source generator of reproducible surface-code quantum error correction datasets for benchmarking decoders.

An open-source generator of surface-code quantum error correction datasets for benchmarking decoders, each file reproducible from its manifest of seeds, versions and a content hash.

Aug–Sep 2026 · shipped

Decoders are only as trustworthy as the data they're scored on. qecgen generates surface-code datasets anyone can regenerate exactly. It grew out of my research internship at Turing BioSciences.

Key facts

  • Stim + PyMatching pipeline
  • Content-hashed, reproducible manifests
  • 583 structural + 8 statistical tests
  • Six-lesson course on top

How it works

  • Pipeline: Stim circuits → detection events → detector error model → sparse check and logical matrices.
  • Exports HDF5, Parquet, NPZ, JSONL or CSV, mapping detection events to logical flips.
  • Validated, with a manifest on every dataset: seeds, pinned library versions, the git commit and a content hash.
  • Realistic noise: per-qubit and per-operation rates, drift, leakage and burst errors.

Tooling

  • Typer CLI: generation, multi-environment runs, drift, scoring and validation.
  • Threshold sweeps with crossing and exponential-suppression fits.
  • FastAPI + React UI: live lattice preview with a cost estimate, threshold charts, a dataset registry.
  • qecgen-learn: a six-lesson course that uses only numbers computed from real runs.

Stack

  • python
  • stim
  • pymatching
  • sinter
  • numpy
  • h5py
  • pyarrow
  • typer
  • fastapi
  • react

Links

#quantum #research #open source

© Haider Ali · London, UK · contact@haiders.website · this is the plain edition of haiders.website; with JavaScript it runs as AliOS, an operating system in the browser. Summary for language models: llms.txt (full text: llms-full.txt).