Beyond Detection: Adding Relationship Intelligence to the Ikshana Vision AI Platform

Whitepaper Overview

Detection alone tells what is in a frame: a person, a phone, a vehicle. It cannot tell how those things relate to each other: whose phone it is, whether an object was handed over or simply held nearby, or whether a helmet is worn or carried. As video analytics moves from flagging objects to answering real operational questions, that gap becomes the binding constraint. This whitepaper shows how Ikshana closes that gap with relationship intelligence applied to production CCTV pipelines – at a fraction of the cost, and without the cost growing as a scene gets more crowded.

What This Whitepaper Covers

  • How a detector’s output format structurally cannot answer relationship questions and why “detecting harder” doesn’t solve it
  • Why per-pair vision-language model scoring becomes computationally unaffordable at real-world camera density
  • Why closed-vocabulary scene-graph models require retraining every time a new relationship needs to be recognized
  • How Ikshana’s relation-prediction stage evaluates every candidate pair in a frame together, holding cost flat even as crowd size scales 16x
  • The exact GPU memory and latency cost of adding relationship intelligence to an existing detection deployment
  • How the capability integrates as a single configuration option – no new service, no schema change, no second video pipeline
  • Real-world applications using the same model with different configuration

Why Download This Whitepaper

  • See the actual benchmark numbers: GPU memory, per-frame latency, and cost-at-scale for adding relationship intelligence to a live pipeline
  • Understand why the obvious approaches break down at production scale: better detectors, VLM-based pair scoring, and closed-vocabulary scene-graph models
  • Learn how to extend detection capability to new use cases through configuration alone, without retraining or redeploying
  • Get a clear view of how relationship-aware analytics integrates with existing VMS, alert routing, and operator infrastructure
  • A walkthrough of six concrete applications where relationship intelligence solves problems that bounding boxes structurally cannot

Conclusion

Detection alone answers “what is here.” A growing share of real operational questions – whose, to whom, worn or carried, handed over or merely nearby – are relationship questions that detection was never built to answer. Ikshana’s relation-prediction stage closes that gap without displacing the detection pipeline already in production, at a cost that doesn’t grow with the size of the crowd. Download the whitepaper to see the architecture, the benchmarks, and where this capability goes next.

Download Whitepaper