Where the model runs is a product decision.
Latency, connectivity, data movement and operating constraints shape useful AI.
Choose around the operating environment.
Running inference locally can reduce network dependence and data movement. Centralising it can simplify updates and make larger compute resources available. Neither choice is universally better.
Start with the actual workload: how quickly the result is needed, how much data is produced, what connectivity can be assumed and who maintains the system.
Keep the interface close to the task.
The most capable model is not always the best fit for a constrained device. The complete system may benefit more from a responsive interaction, predictable memory use or a clear fallback than from a small improvement on an isolated benchmark.
In video analytics, the path includes capture, preprocessing, inference, tracking, transport and presentation. Measure that path rather than treating model execution time as the whole latency budget.
Design for the day the connection drops.
An event-day scoring tool and a camera-analysis system have different workloads, but both need to explain what happens when part of the system is unavailable. People need to know what was saved, what is still local and what has reached the shared record.
A fallback is part of the experience. Plan it alongside the normal path, test it under the conditions people will encounter, and make its limits visible without asking the operator to diagnose the infrastructure.