MIT Eyes turns the CCTV you already own into a tireless AI analyst — detecting people, vehicles, animals, fire, smoke and weapons in real time, recognising faces and plates, making every second of footage searchable, and emailing the right person the moment something matters. All of it self-hosted, on your hardware.
RTSP, ONVIF, MJPEG, HLS, RTMP and WebRTC — IP cameras, NVRs, drones or phones stream straight in. No proprietary hardware, ever.
The built-in network scanner sweeps a subnet, lists every live device and infers which ones are cameras — so onboarding a site you inherited doesn't start with a spreadsheet.
Author what to detect and recognise once — “Perimeter, high accuracy”, “Lobby, faces only” — then assign that profile to any number of cameras. Retune the whole estate by editing one schema.
Draw the part of the frame that matters and the pipeline ignores the rest. The street behind your gate stops waking the AI — and stops generating noise.
People, vehicles and animals detected live with motion-gated AI, so compute is spent on moments that matter — not on empty hallways.
Fire, smoke, sparks, weapons and knives trigger instant on-screen and email alerts — seconds matter, and MIT Eyes reacts in under one.
Every face is embedded by two independent recognisers (ArcFace + AdaFace). Hard angles, low light, ageing footage — one engine catches what the other misses.
ANPR with fuzzy matching that forgives single-character OCR misreads, plus a dedicated tracking mode that votes every frame of a car's pass into one plate.
Drag camera, subject, time, frequency and email onto a canvas and you have an alert rule. “Unknown face at the back door after 20:00, at most one email an hour” takes a minute to build and no developer.
A security-operations command view: system posture at a glance, live threat feed, camera health, activity per hour and a watchlist of what still needs a human.
Pick the accuracy/speed tier per function — detection, faces, plates, hazards. Weights download in the background and the switch only flips once they're ready, so nothing ever goes dark.
Point MIT Eyes at years of recorded footage — local disks, network shares or cloud buckets — and retire the originals as a searchable database of detections.
“Show me every white van since Tuesday” is a two-second query, not a two-day shift of watching timelines.
Model the real world — sites, buildings, camera locations — so every detection carries where it happened and multi-site estates stay legible.
Traffic patterns, occupancy trends and hardware sizing built in — evidence for decisions, not gut feel.
Runs entirely on your own servers, with or without a GPU. No cloud account, no per-camera subscription, no footage leaving the building.
| The traditional way | MIT Eyes | |
|---|---|---|
| Finding an incident | ✕Hours of scrubbing timelines, hoping you don't blink at the wrong moment | ✓Type a query — every person, plate and event is indexed and found in seconds |
| When you learn about it | ✕After the damage is done, during footage review | ✓The moment it happens — real-time alerts for hazards, faces and plates |
| Who tells you | ✕Someone has to be watching the video wall — and be the right someone | ✓Rules you draw once email the right person, with the evidence attached |
| Who's watching | ✕A human operator whose attention collapses after ~20 minutes | ✓AI on every frame of every camera, around the clock, with equal focus |
| Retuning a camera | ✕A vendor visit, a per-device config, and a change window | ✓Edit one vision schema — every camera assigned to it follows within seconds |
| False alarms | ✕Every passing car on the road behind the fence | ✓Detection areas confine the AI to the ground you actually own |
| Storage | ✕Terabytes of raw 24/7 footage nobody will ever watch | ✓Meaningful, indexed events with photo evidence — a fraction of the footprint |
| Hardware | ✕Proprietary NVRs, licensed channels, forklift upgrades | ✓Runs on your existing cameras and commodity servers |
| Your data | ✕Locked in a vendor cloud, per-camera subscription fees | ✓Fully self-hosted on your premises — your footage never leaves the building |
Every face is stored with both ArcFace and AdaFace embeddings. Switch engines per camera at any time — no re-enrolment, no lost history. Almost nobody else in the market does this.
Vision schemas turn detection settings into a reusable profile. Change one schema and every camera carrying it retunes itself — the difference between managing 8 cameras and managing 800.
The workflow builder turns “tell facilities when an unknown face appears at the loading bay after hours” into a five-node diagram. Competitors ship an API and a quote for professional services.
Motion gates and drawn detection areas discard irrelevant frames before the heavy AI runs, so one modest GPU covers a whole site — competitors quote a server rack.
Faces and plates are sensitive data. MIT Eyes keeps them on your hardware with zero cloud dependency and zero recurring per-camera fees.
Years of DVR recordings stop being dead weight: mine them into searchable detections, then reclaim the storage.
The same cameras and the same install power MIT Parking, MIT Attendance and MIT Traffic — each added module multiplies the return on the first.
Competitors sell four systems: four installs, four vendors, four invoices. MIT sells one platform where every module reuses the cameras, the AI and the data of the others — so each solution you add costs a fraction and returns in full.
See MIT Eyes running on your own cameras — a pilot takes a day, not a quarter.