Vision
Vision · cuthub-vision-1
Footage understanding built for editors. Not on sale yet: request a trial.
Four levels
| Level | What you get |
|---|---|
| probe | Stream fingerprint, GOP, I-frames and scene cuts, static segments, duplicate frames, motion energy. |
| shots | Frame-accurate shot boundaries (hard cut, fade, dissolve, wipe, flash), with four frames reviewed around every cut. |
| cards | A card per shot: motion, camera, tonality, sound class, transcript, and a strict-JSON visual description, plus a keyframe contact sheet. |
| look | A per-second full-frame sheet with a manifest, made for an agent to look at. |
Frames are the ground truth
Model-reported seconds and brand names are not trusted; everything is checked against frames. Accuracy numbers will be published here, in the same expectation, verdict, evidence format as Sansi, once the ground-truth set is reviewed. We publish nothing until then.
Privacy
The cards level can send keyframes to a third-party vision model; probe, shots, and look do not leave our servers. A mode with no third-party model will be available for footage that must not leave.
Request a trial
Email us with what footage you work with and roughly how many minutes a month.