This decision is usually made by preference and then justified afterwards. It is worth making it on three numbers instead: how fast the answer is needed, how much bandwidth you have, and where the data is legally allowed to sit.
Latency: what has to happen before the moment passes
If the output triggers a physical action — a barrier, a reject arm, a tunnel alarm — the round trip to a cloud region is usually too slow and too dependent on a link you do not control. If the output feeds a dashboard someone reads hourly, latency is irrelevant and the cloud is easier to operate.
Bandwidth: streaming video is expensive forever
A single 1080p stream at a usable frame rate runs a few megabits per second, continuously. Multiply by camera count and by every month of the contract. Edge processing inverts this: frames stay local, and only events and requested clips travel.
- Central processing: continuous upload of every stream, every hour
- Edge processing: events and thumbnails, with clips pulled on demand
- Hybrid: edge detection, cloud for storage of confirmed events only
Residency: where the data is allowed to be
Some footage cannot leave the site, the state or the country. Once that constraint exists the discussion is over — process locally and send only what is permitted. Establish this before designing anything, not after.
A short decision table
- Sub-second action required → edge, without exception
- Thin or unreliable link → edge, with event-only upload
- Data cannot leave the premises → edge or on-premise server
- Many sites, no real-time action, good links → central or cloud
- Heavy retraining and analysis workloads → cloud, with edge inference
The common answer in practice is hybrid: inference at the edge because that is where latency and bandwidth force it, training and long-term analysis centrally because that is where the compute is worth paying for.
What the edge costs you
It is not free. Edge means hardware at every site, a way to update models remotely, monitoring so a dead box is noticed, and physical access when something fails. Budget for that operational load rather than discovering it in year two.