Experience · 03
Turing BioSciences — quantum error correction
By Haider Ali · Quantum & Machine Learning Research Intern, Turing BioSciences ·
Haider Ali was a Quantum & ML Research Intern at Turing BioSciences in 2026: he benchmarked QEC decoders on real device data and built qecgen.
Quantum & ML Research Intern. Benchmarked quantum error-correction decoders on real device data, built QEC data pipelines as open-source qecgen, and curated biotech sequence data.
Jun–Sep 2026 · quantum & ml research intern · shipped

A remote internship reporting directly to the founder, across Turing BioSciences' quantum and biotech workstreams: benchmarking quantum error correction decoders on real device data, and curating sequence data for biotech ML.
Key facts
- ML vs classical QEC decoders
- Real quantum-device data
- qecgen: 583 structural + 8 statistical tests
- Remote · reported to the founder
Quantum error correction
- Decoders: developed and benchmarked machine-learning decoders against established classical baselines.
- Real hardware: prepared quantum-computer data alongside simulation, to test decoders under realistic noise.
- Pipelines: Python generators of large-scale QEC datasets with known ground truth and engineered features.
qecgen (open source)
- Generator: Stim circuits → detection events → HDF5, Parquet, NPZ or JSONL.
- Reproducible: every file carries a manifest of seeds, versions and a content hash.
- Tooling: a CLI for generation, threshold sweeps, drift and scoring, with a FastAPI and React UI.
- qecgen-learn: a six-lesson crash course where every number comes from a real run.
Biotech data
- FASTA datasets curated and preprocessed for vaccine and drug development and mutation prediction.
- Provenance-first pipeline: public genomic data to model-ready data; it refuses to invent missing labels.
Stack
- python
- stim
- pymatching
- sinter
- numpy
- h5py
- pyarrow
- fastapi
- react
- typescript
Links
#quantum #ml #research #data