In today's world, the most important component of video surveillance is AI. The vendors see AI as a game-changer that can identify threats in real time, minimize false alarms, and optimize operational efficiency across various sectors. The proposition on paper makes a lot of sense.
It is generally a good performer in the controlled demonstrations. However, that's not so in the real world. Most failures will start with the gap between demo performance and operational reality.
A common and troubling challenge with organizations using AI video monitoring systems is missed events. But regularly.
These failures are not frequently reflected in the performance dashboard since most systems report detections, not missed outcomes. Rather, they re-emerge in the wake of an incident when teams find out the system didn't work or didn't work quickly enough to respond or that it created too much noise and the signal was lost.
To grasp why this is happening, it's important to peel back the layers of the “AI accuracy” story and explore what happens in the real world. When events are missed, it is not due to a single failure point.
They are the result of the various constraints faced in data quality, environmental variability, infrastructure constraints and human workflows. In practice, this is not a problem to solve in a model – it is a system design problem.
Why AI Video Monitoring Systems Miss Events?
Environmental variability, mismatch in training data, noise (too many false positives), hardware constraints, processing time, and context are all sources of events being missed in AI video monitoring systems. These are not stand-alone factors.
They can propagate throughout the system and result in delayed alerts, suppressed detections or no detection at all of events of significance. The important point is this: Better model accuracy is not the answer to missed events. Reliability is a function of the behaviour of the whole system when it is exposed to the real world.
Detection Without Understanding: The Fundamental Limitation of AI Surveillance
Object detection, which refers to the ability of a video monitoring system to recognize objects like people, vehicles or things in a frame, is the foundation of most AI video monitoring systems. This capability has come a long way, but it's just a fraction of what is needed in real-world monitoring. In practice an organisation is not interested in objects. They are interested in events and they observe what happens.
It is not about a person but about what they are doing and where they are doing it and how they are changing their behaviour. For instance, a standing person who behaves normally for a few seconds, but who stands at a specific spot for a long time or enters and exits from restricted areas repeatedly is suspicious.
This creates a crucial gap between the detection and interpretation. Most systems are frame-by-frame processing which do not model temporal sequences or behavioural context. Consequently, they will produce many detections but they will not be able to distinguish whether these detections are significant or not. This results in missed events and over alerts, which means that the system is technically alive but operationally dead.
Environmental Complexity: Why Real-World Conditions Break AI Models
AI models are very sensitive to the conditions under which they are trained. The training data has some flexibility, but it seldom represents the entire range of real-world environments. During deployment, systems are subjected to a myriad of variables that have a large impact on visual input.
One of the most influential factors is lighting conditions. Differences in light intensity due to the time of day, weather, or artificial light sources impact contrasts, colour uniformity and how shadows are created. In low light, there is noise in the sensor which can make it harder to see the features that models use to locate, causing problems with detection.
The weather further exacerbates the problem. Rain, fog, dust and glare affect the visual scene, making it less visible and changing the boundaries of the objects. These effects pose a problem especially to models based on edge detection or texture recognition.
But the level of complexity of the scene is also significant. In a crowded environment objects often overlap, move in an arbitrary manner and are partially occluded. The background movement, including swaying trees, working machinery or moving reflections, adds to the variability and may be mistaken for true movement.
These factors combine to not only impact accuracy but also the consistency of operation. Reliability can be a problem, as systems might work well in one circumstance and not in another.
False Positives and Noise: The Leading Cause of Operational Failure
Over-detection is one of the most important factors for missed events. AI systems are meant to detect changes in a scene; however, not all changes are relevant in a real environment.
Noise is alerts that are generated by actions that are not actionable. These may be due to insects crawling on the lens, animals crossing the frame, wind blowing over foliage or debris, or light changes from reflections or passing vehicles. In low light, the noise from the sensor can cause false detections.
This noise can produce thousands of alerts a day in large-scale deployments. Every alert needs to be addressed, even if it is to be dismissed. This imposes a heavy workload on the monitoring staff who have to check the incoming signals constantly to decide whether they have to handle them or not.
It's not just about inefficiency. If operators are inundated by false positives, they will experience alert fatigue, which is a decrease in sensitivity to alerts when there are many and they are not very valuable. The more fatigued something is, the longer it takes to respond, the less attention it gets, and the more likely it is that it will miss out on real events.
This results in the paradoxical situation that the more activity systems can detect, the less activity systems can be detected because they overload the human operator.
The Human Factor: How Alert Fatigue Leads to Missed Incidents
In real monitoring situations, the human operator plays an important role in the decision-making process. For critical tasks, even the most sophisticated AI systems require human validation.
If volumes of alerts are reasonably small, operators can look at each event and make sure that it is handled properly. But in a noisy environment, the system becomes a source of information overload, rather than a decision support system.
A brief assessment of each alert is needed, typically 10-20 seconds. On a large scale, this equates to low-value, repetitive hours. This will result in shorter attention span, slower reaction times, and greater heuristic over analytic reliance over time.
The result isn't only inefficiency, it's risk. With real events, it's more likely that the event is postponed or completely missed. In this context, missed events are not just technical mistakes but a product of human and system interaction when under stress.
Training Data Limitations: Why AI Fails Outside Controlled Scenarios
The data that trains the AI models is the foundation of the models. But there is typically a wide mismatch between the training data environment and the deployment environment.
Training datasets are usually carefully handcrafted with good examples of objects and events, with consistent labelling and controlled variation. This will help the model to perform better in the development phase, but it will restrict the generalisation capabilities of the model.
When deploying models, you are likely to come across scenarios that are not represented during training. These include rare edge cases, region-specific behaviors and unusual camera angles. This results in a reduction of confidence scores and consequently, missed detections or incorrect classifications.
No simple incremental solutions to this problem exist. It requires ongoing data gathering and retraining to real-world conditions, which many systems don't have.
Hardware and Video Quality: The Foundation of AI Performance
While AI performance is usually assessed on its own, in reality there's no separating the two. Video feed quality is important to the quality of the analysis.
Poor low-light performance in cameras results in noisy images that make it difficult to see what's important. Low dynamic range means that bright parts and dark parts of a scene can't be captured in a single shot. Wrong lens choice can make distant objects appear less clear or change the perspective.
The positioning of the camera is also crucial. No algorithm can fix gaps that occur in visibility due to blind spots, improper angles, and insufficient coverage.
AI isn't producing new info. It is able to find patterns in data. The input is not clear; the output will not be reliable.
Latency and Processing Architecture: When Timing Becomes a Failure Point
In time sensitive applications, latency becomes as crucial as accuracy. Many of the stages, such as video transmission, video frame processing, model inference, and alert delivery can experience delays.
Cloud-based systems add network-layer dependencies, which can lead to latency issues, especially in areas with low bandwidth or unreliable network connections. Edge-based systems, which aim to minimize latency, lack the computational power to support more sophisticated models or higher frequencies of model processing.
Such compromises are generally not obvious to end users but they do affect performance. An alert received 2 or more seconds after an incident has occurred can be technically correct, but operationally irrelevant.
Frame Sampling and Temporal Gaps in Detection
To slow down the calculation, many systems process only a fraction of the video frames. This way, the efficiency is increased, but gaps in temporal coverage are introduced.
Any events happening between processed frames are essentially ignored by the system. This can be especially troublesome when you're dealing with a short-duration or high-speed event and you want to start and/or end within a fraction of a second.
For many deployments, frame sampling is a practical requirement; however, it is important to carefully manage it to prevent detection blind spots.
Model Drift: The Gradual Decline of System Accuracy
System performance degrades over time even if the system is performing well at first. The data distribution moves as environments change, and the model is no longer as aligned as it should be to current conditions.
This can be because of changes in layout, lighting or behaviour patterns. If the model is not updated frequently, its performance will deteriorate over time and result in more missed events and false positives.
Drift is especially difficult due to the fact that it is incremental. There's no one critical point of failure, just a gradual loss of reliability.
Why Missed Events Are a System-Level Problem
The common thread in all these factors is that no one factor is responsible for missed events. They are a product of several interacting parts that are not completely in sync.
Model accuracy improvements alone cannot solve issues of environmental variability, hardware limitations or operational constraints. Likewise, infrastructure optimization doesn't solve data quality or event interpretation problems.
This means a system-level approach, connecting data, hardware, processing and human workflows together to create an effective AI video monitoring.
Building AI Video Monitoring Systems That Don’t Miss What Matters!
The future of AI video monitoring is not just about refining detection models; it's about rethinking the design and deployment of these systems.
Organisations require systems that can be built for real environments, that can tolerate variability, that can minimise noise, that can evolve as time goes on, and that can present meaningful information to the operators without swamping them.
That is where a development such as Spotem AI video intelligence platform comes in more advanced. Unlike point solutions, Spotem is an AI security monitoring system, considered as an integrated system where data quality, data processing architecture, behavioural modelling and alert relevance are aligned.
AI CCTV cameras with real-time alerts provide more precise results, reduced overheads and quicker decision making as they tackle the underlying causes of missed events, not just detection metrics.
This is very important in high-stakes situations. The best measure of the worth of AI video monitoring is not what it finds. It is determined by its effectiveness in ensuring that nothing important is missed.
Top comments (0)