UNICORN WHO DEV
← Back to catalogue
Datasets / fireviewer

v2.5-ready-to-train

Public dataset for object detection, documented on the Hub.

Open the public source
Overview

Public dataset for object detection, documented on the Hub.

Downloads
0
Updated
Oct 03, 26
Size
10K–100K
Source
Hugging Face
Status
Connected source
Public dossier

Documentation and scope

Structured content from the public source documentation.

01

Public documentation

38,965 unique native images after merging FireViewer detection V2 and the published negative corpus, including its 195 very difficult JPEG candidates before deduplication.

Splits: {"train": 31657, "validation": 3619, "test": 3689}. The 80/10/10 target is adjusted to keep groups intact. Seed: 42.

Classes: 1 smokevisible, 2 flamevisible. Boxes are COCO xywh in original raw pixel coordinates. Negatives have zero annotations. Source points remain provenance only.

Visual referencev2.5-ready-to-train · Object Detection
02

Train

Run the included preparecoco.py with --revision IMMUTABLECOMMIT --output /content/v25. It materializes train, valid and test, each with native images and annotations.coco.json. Use valid for selection and keep test closed until final evaluation. Models map category IDs 1/2 to internal labels 0/1.

Decode original pixels without automatic EXIF rotation. No lossy recompression or image resizing was used to build this dataset.

03

Split integrity and limits

Historical model-exposure cases and their detected related groups are kept in train. Validation/test images are disjoint by native and decoded-pixel identity, normalized source origin, known capture groups and the documented perceptual checks.

The split receipt describes the algorithm and counts; the exclusions audit records duplicates and conflicting identical-image annotations. This is not a claim of exhaustive scene identity or individual human validation of every annotation. Original annotation authority is retained. Native negative difficulty is unrated; very difficult JPEGs are separate.

04

Sources and licences

Source datasets: fireviewer/fire-smoke-detection-corpus-v2 at c1336415e1f9e7f67f332a6ec212345fdf3ffdb2 and fireviewer/fireandsmokedetectionveryhardnegative at 04ed0590954245b524dd265a6da63462b2e50f62. Mixed source-specific licences apply; this release does not grant new image rights. Consult provenance, per-image source references and the original sources.

All published Parquet files were downloaded again, reloaded, and checked against native image hashes, final objects and split membership before public release.

05

Materialization dependencies

Install requirements.txt in a remote Python environment, then run preparecoco.py against an immutable revision. The materializer uses eight parallel downloads and verifies every native image before creating COCO folders.

python -m pip install -r requirements.txt python preparecoco.py --revision --output /content/v25

Native negatives retained: 12,653. Very difficult negatives retained: 195 (153 train, 21 validation, 21 test). The final corpus contains 50,591 boxes. See audit/release-verification.json and audit/metadata-verification.json for verification scope and revision identities.

Provenance

A record traceable to its source

Editorial information, metadata and documentation remain connected to the original public repository.

Public identifier
fireviewer/v2.5-ready-to-train
Latest activity
October 3, 2026
Read the complete documentation
Technologies
image