Skip to content

Vision

Evaluating Sentinel for Dense Crowd and Public-Space Scenes

How to interpret a compact person-detection release, rebuild its metrics on your camera population, and connect detections to a responsible review workflow.

Innomium Vision Team5 min read
Computer vision evaluation across dense public-space camera scenes

Sentinel is a compact person-detection artifact intended for scenes where occlusion, object density, lighting, and runtime limits matter. Its public evaluation snapshot is a starting point for inspection, not a performance guarantee for an airport, stadium, retail floor, or public venue.

A responsible evaluation rebuilds the task definition and metrics on the camera population where the system will operate.

Define what a person detection supports

Counting, queue estimation, occupancy, restricted-area review, and crowd-density awareness require different event logic. Decide whether partial persons count, how zones are defined, and which misses or false alarms create real operating cost.

Design the scene population

Sample entrances, concourses, reflective surfaces, shadows, screens, staff areas, and periods of low and high density. Separate results by camera, elevation, object scale, lighting, and occlusion. Include empty scenes and non-person distractors.

Measure more than boxes

Use precision and recall at the event-relevant operating point, plus latency and stability on target hardware. If the workflow depends on counts or density, evaluate that downstream result directly instead of assuming detection metrics transfer.

Keep the operator in the design

Alerts should include enough visual and temporal context to review quickly. Monitor dismissals and corrections as workflow evidence, while respecting privacy, retention, and access requirements.

Define crowd conditions as measurable slices

Crowd evaluation should separate density, occlusion, perspective, camera height, motion, lighting, and scene type. A single count error hides whether the model is reliable in sparse entrances but weak in dense, distant regions. Create scene-level and region-level labels that match the decision the operator will make.

Public-space imagery requires a clear purpose, retention policy, access boundary, and review process. Prefer aggregate or event-level outputs when identities are unnecessary. Document which inferences are intentionally excluded so later feature requests do not silently broaden surveillance scope.

Evaluate workload as well as detection

Measure false alerts per camera-hour and per operating condition, not only precision on selected frames. Correlated false positives can overwhelm an operator even when an image benchmark appears strong. Replay continuous footage with quiet periods, transitions, weather, and camera vibration.

Thresholds should reflect action capacity and consequence. A warning used to allocate staff can tolerate different uncertainty from an emergency escalation. Monitor alert acknowledgement, dismissal reasons, and missed-event review to determine whether the system improves situational awareness rather than merely producing more signals.

Executive decision record

The decision is which aggregate crowd observations are necessary for the stated public-space purpose and which inferences should remain explicitly out of scope. Write that decision before selecting a model, vendor, framework, or implementation pattern. A written boundary keeps technical exploration connected to the operating outcome and makes it possible to explain why the organization advanced, revised, or stopped the work.

Approval should depend on density and scene slices, continuous camera-hour replay, workload measures, privacy review, retention controls, and operator acceptance. The evidence does not need to remove every uncertainty, but it should address the uncertainty capable of changing value, architecture, risk, or ownership. Record the baseline, assumptions, unresolved questions, and the person accepting the next stage.

Failure boundary and operating ownership

The central failure to guard against is relying on an image-level average that hides correlated alerts, dense-scene misses, or expansion into unnecessary identity inference. Treat that condition as a testable scenario. Define how the system detects it, what users experience, which action is prevented or reversed, and what evidence reaches the person responsible for recovery.

Long-term accountability sits with the public-space operator and privacy authority, with engineering responsible for bounded outputs and measured performance. Supporting specialists can provide platforms, research, review, or delivery capacity, but they cannot substitute for an owner who controls policy and operating change. Name that owner before production and include the ownership path in release evidence and incident procedure.

A practical 90-day application plan

During the first 30 days, convert density and scene slices, continuous camera-hour replay, workload measures, privacy review, retention controls, and operator acceptance into a bounded evidence plan. Assign each artifact to a named contributor, identify the representative inputs required, and agree on the comparison baseline before implementation expands. The objective of this period is to expose the assumption most likely to invalidate the work while the cost of changing direction is still low.

During days 31 through 60, build or instrument the smallest complete workflow that can support the decision about which aggregate crowd observations are necessary for the stated public-space purpose and which inferences should remain explicitly out of scope. Include the real data and authorization path where feasible, record exceptions, and review difficult cases with the people who own the underlying process. Resist adding breadth until the team can explain the measured behavior of this narrow slice.

During days 61 through 90, test the boundary represented by relying on an image-level average that hides correlated alerts, dense-scene misses, or expansion into unnecessary identity inference. Exercise degraded dependencies, ambiguous inputs, recovery, and handoff rather than demonstrating only successful cases. End the period with a written advance, revise, or stop decision that cites evidence, residual exposure, expected operating cost, and the next authority boundary.

The review should be accepted by the public-space operator and privacy authority, with engineering responsible for bounded outputs and measured performance. That group should confirm not only that the system can work, but that ownership, support capacity, monitoring, and change control are credible. If those conditions are absent, the responsible outcome is another bounded learning stage rather than an unsupported production commitment.

Practical checklist

  • Write the event definition and zone rules.
  • Sample dense, sparse, empty, and difficult scenes.
  • Measure by camera and operating condition.
  • Profile the exported model on target hardware.
  • Evaluate downstream counts or events directly.
  • Design review, retention, and audit behavior.

Continue reading

  • [Edge vision evaluation protocol](/edge-vision-evaluation-protocol)
  • [Building the Innomium Vision Layer](/building-the-innomium-vision-layer)

Sources and further reading

  • [ONNX Runtime edge deployment](https://onnxruntime.ai/docs/tutorials/iot-edge/)
  • [NIST AI RMF](https://www.nist.gov/itl/ai-risk-management-framework)

Want production AI shipped with the same discipline?

Talk with Innomium about vision models, long-context systems, or a focused engineering program.