UNICORN WHO DEV
← Retour au catalogue
Datasets / fireviewer

v2.5-ready-to-train

Dataset public pour object detection, documenté sur le Hub.

Ouvrir la source publique
Vue d’ensemble

Dataset public pour object detection, documenté sur le Hub.

Téléchargements
0
Mis à jour
03 oct. 26
Taille
10K–100K
Source
Hugging Face
Status
Source connectée
Dossier public

Documentation et périmètre

Contenu structuré depuis la documentation publique de la source.

01

Documentation publique

38,965 unique native images after merging FireViewer detection V2 and the published negative corpus, including its 195 very difficult JPEG candidates before deduplication.

Splits: {"train": 31657, "validation": 3619, "test": 3689}. The 80/10/10 target is adjusted to keep groups intact. Seed: 42.

Classes: 1 smokevisible, 2 flamevisible. Boxes are COCO xywh in original raw pixel coordinates. Negatives have zero annotations. Source points remain provenance only.

Repère visuelv2.5-ready-to-train · Object Detection
02

Train

Run the included preparecoco.py with --revision IMMUTABLECOMMIT --output /content/v25. It materializes train, valid and test, each with native images and annotations.coco.json. Use valid for selection and keep test closed until final evaluation. Models map category IDs 1/2 to internal labels 0/1.

Decode original pixels without automatic EXIF rotation. No lossy recompression or image resizing was used to build this dataset.

03

Split integrity and limits

Historical model-exposure cases and their detected related groups are kept in train. Validation/test images are disjoint by native and decoded-pixel identity, normalized source origin, known capture groups and the documented perceptual checks.

The split receipt describes the algorithm and counts; the exclusions audit records duplicates and conflicting identical-image annotations. This is not a claim of exhaustive scene identity or individual human validation of every annotation. Original annotation authority is retained. Native negative difficulty is unrated; very difficult JPEGs are separate.

04

Sources and licences

Source datasets: fireviewer/fire-smoke-detection-corpus-v2 at c1336415e1f9e7f67f332a6ec212345fdf3ffdb2 and fireviewer/fireandsmokedetectionveryhardnegative at 04ed0590954245b524dd265a6da63462b2e50f62. Mixed source-specific licences apply; this release does not grant new image rights. Consult provenance, per-image source references and the original sources.

All published Parquet files were downloaded again, reloaded, and checked against native image hashes, final objects and split membership before public release.

05

Materialization dependencies

Install requirements.txt in a remote Python environment, then run preparecoco.py against an immutable revision. The materializer uses eight parallel downloads and verifies every native image before creating COCO folders.

python -m pip install -r requirements.txt python preparecoco.py --revision --output /content/v25

Native negatives retained: 12,653. Very difficult negatives retained: 195 (153 train, 21 validation, 21 test). The final corpus contains 50,591 boxes. See audit/release-verification.json and audit/metadata-verification.json for verification scope and revision identities.

Provenance

Une fiche traçable jusqu’à sa source

Les informations éditoriales, les métadonnées et la documentation restent reliées au dépôt public d’origine.

Identifiant public
fireviewer/v2.5-ready-to-train
Dernière activité
3 octobre 2026
Consulter la documentation complète
Technologies
image