Projects · 02
qecgen — datasets for quantum decoders
By Haider Ali ·
qecgen, by Haider Ali, is an open-source generator of reproducible surface-code quantum error correction datasets for benchmarking decoders.
An open-source generator of surface-code quantum error correction datasets for benchmarking decoders, each file reproducible from its manifest of seeds, versions and a content hash.
Aug–Sep 2026 · shipped
Decoders are only as trustworthy as the data they're scored on. qecgen generates surface-code datasets anyone can regenerate exactly. It grew out of my research internship at Turing BioSciences.
Key facts
- Stim + PyMatching pipeline
- Content-hashed, reproducible manifests
- 583 structural + 8 statistical tests
- Six-lesson course on top
How it works
- Pipeline: Stim circuits → detection events → detector error model → sparse check and logical matrices.
- Exports HDF5, Parquet, NPZ, JSONL or CSV, mapping detection events to logical flips.
- Validated, with a manifest on every dataset: seeds, pinned library versions, the git commit and a content hash.
- Realistic noise: per-qubit and per-operation rates, drift, leakage and burst errors.
Tooling
- Typer CLI: generation, multi-environment runs, drift, scoring and validation.
- Threshold sweeps with crossing and exponential-suppression fits.
- FastAPI + React UI: live lattice preview with a cost estimate, threshold charts, a dataset registry.
- qecgen-learn: a six-lesson course that uses only numbers computed from real runs.
Stack
- python
- stim
- pymatching
- sinter
- numpy
- h5py
- pyarrow
- typer
- fastapi
- react
Links
#quantum #research #open source