A drop-in replacement for pycocotools that computes COCO-style AP/AR for
object detection, instance segmentation and keypoint detection. Ground-truth
indexing, IoU/OKS matching and precision–recall accumulation run in a
multithreaded Rust core (Rayon), exposed to Python through PyO3.
- Bit-exact metrics. The complete
precision,recallandscoresarrays are bit-identical to pycocotools on every tested input, not only the rounded AP/AR summary. - Fast. On COCO val2017, 18–54× lower wall-clock time than pycocotools, 7–9× lower than faster-coco-eval and 1.8–2.6× lower than hotcoco.
- Memory-efficient. 58–86% lower peak RSS than pycocotools, and the lowest of all four evaluators on every task.
- Drop-in. The same
COCO/COCOevalAPI forbbox,segmandkeypoints, plus the LVIS federated protocol. Prebuilt wheels for Linux, macOS and Windows (CPython 3.8–3.14).
Status: alpha. Before replacing the reference evaluator in your pipeline,
check parity on your own evaluation parameters and COCOeval subclasses.
pip install ultrafast-pycocotoolsPrebuilt wheels need no Rust compiler; NumPy is installed as a dependency.
Other platforms build from source with a stable Rust toolchain:
pip install git+https://github.com/developer0hye/ultrafast-pycocotools.
from ultrafast_pycocotools import COCO, COCOeval
gt = COCO("instances_val2017.json") # ground-truth annotations
dt = gt.loadRes("detections.json") # detection results in COCO format
evaluator = COCOeval(gt, dt, "bbox") # or "segm", "keypoints"
evaluator.run() # evaluate() + accumulate() + summarize()
print(evaluator.stats_as_dict) # AP, AP50, AP75, APs, APm, APl, AR@1, ...For LVIS, pass lvis_style=True to apply the federated annotation protocol
(negative and not-exhaustive category lists, 300 detections per image) and the
official metric names (AP, APr, APc, APf, AR@300, ...). See the
LVIS guide.
If a training or validation framework imports pycocotools internally,
register the replacement before importing that framework:
from ultrafast_pycocotools import init_as_pycocotools
init_as_pycocotools() # `import pycocotools` now resolves to ultrafastThis patches sys.modules for the whole Python process.
COCO val2017 (all 5,000 images) with cached YOLO26n, YOLO26n-seg and YOLO26n-pose detections: 733,070 boxes, 724,953 instance masks and 134,663 pose instances. Model inference is excluded; the timed region covers JSON parsing, ground-truth indexing, matching, accumulation and summarization.
| Task | pycocotools 2.0.11 | faster-coco-eval 1.8.0 | hotcoco 1.0.1 | ultrafast 0.1.11 |
|---|---|---|---|---|
| bbox | 47.55 s · 1,925 MB | 7.51 s · 1,781 MB | 2.33 s · 2,307 MB | 0.89 s · 272 MB |
| segm | 49.38 s · 2,191 MB | 15.94 s · 2,629 MB | 5.30 s · 3,403 MB | 2.33 s · 838 MB |
| keypoints | 9.83 s · 603 MB | 4.54 s · 603 MB | 0.99 s · 642 MB | 0.54 s · 256 MB |
| Bit-identical to pycocotools | reference | ✗ (precision ≤ 2.2e-16 off on bbox/segm) |
✗ (scores differ on bbox/segm) |
✓ all tasks |
Wall-clock time · peak RSS; median of 6 runs, each in a fresh process, on an Intel Core i5-10400 with a 2-thread pool (pycocotools is single-threaded), measured 2026-09-24. Full report, method and raw data · All benchmarks, other hosts and historical results
- RF-DETR: optional
ufcocobackend for bbox and mask mAP during training and validation, merged in roboflow/rf-detr#1449. Benchmark
Proposed integrations under review: Ultralytics (validation) · torchvision · TorchMetrics · SAHI · SAM 3 · RT-DETR · DEIMv2
The public COCO, COCOeval and mask (RLE) APIs are tested against
pycocotools, covering annotation query order, loadRes, compressed and
uncompressed RLE, custom Params (IoU thresholds, area ranges, maxDets),
iscrowd handling, score ties and subclass overrides. A few internals
intentionally differ:
- Per-image matching records (
evalImgs) are not stored by default; passstore_eval_imgs=Trueif your code reads them. - Box-only detections omit the polygon
segmentationderived from each box; passderive_segmentation=Trueif you read that field. COCO(annotation_dict)borrows the dictionary instead of deep-copying it, and ground-truth annotations are not rewritten in place.
Bit-exactness claims cover the tested inputs and configurations. See the compatibility notes.
The same matching results provide per-category AP, precision–recall curves, per-detection TP/FP matches and a confusion matrix. Boundary IoU is available as an additional evaluation mode (an extension, not a pycocotools metric):
evaluator.per_category_stats()
evaluator.pr_curve(cat_id=1, iou_thr=0.5)
evaluator.matches(iou_thr=0.5)
evaluator.confusion_matrix()See the implementation notes.
git clone https://github.com/developer0hye/ultrafast-pycocotools.git
cd ultrafast-pycocotools
pip install -e ".[test,lvis-test]"
python bench/fetch_lvis_fixture.py
python -m pytest -q
cargo test -p ufcoco-corepython bench/reproduce.py quick --out bench/out/quick --verify-published quick
runs an end-to-end parity check on synthetic data, with no dataset download
or model weights. See the
reproduction guide,
CI and
release process.
Documentation, comments and examples are written in English.
BSD-2-Clause. The evaluation protocol and API follow pycocotools by Piotr Dollár and Tsung-Yi Lin (BSD-2-Clause). LVIS verification uses the official LVIS API. Design decisions are in DESIGN.md.