Conceptio › Archive › arXiv CS
arXiv CSopen access

On-the-Fly Homographies Calibration for Multi-Camera Tracking

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

On-the-Fly Homographies Calibration for Multi-Camera Tracking David Voihanski and Mor Sinai

Ben Zion Bobrovsky

Juganu Or Yehuda, Israel {davidv, mors}@juganu.com

Tel Aviv University Israel [email protected]

arXiv:2609.18582v1 [cs.CV] 16 Sep 2026

© 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.

Abstract—Precise multi-camera tracking traditionally relies on rigorous 3D site calibration, yet this requirement is often operationally impossible in large-scale deployments. Privacy regulations frequently prohibit recording video for offline calibration; limited bandwidth precludes synchronizing high-resolution streams from hundreds of cameras; and covering immense physical sites with calibration targets is logistically infeasible. We present a multi-camera homography calibration system designed to overcome these barriers through ”on-the-fly” geometric refinement. Starting from coarse manual homographies, we introduce a centroid-based projection optimization (PO) that continuously aligns the ground-plane geometry using live detection streams. Because PO operates asynchronously on alreadytransmitted, lightweight metadata, it adds zero computational latency to the real-time tracker. This allows the system to adapt automatically to camera movements or environmental changes without human intervention. This optimized geometry feeds a multi-camera bird’s-eye-view (BEV) tracker that fuses detections and unifies trajectories across zones. Crucially, by operating strictly on live anonymous metadata, our solution ensures a privacy-safe, zero-overhead, and resilient tracking pipeline that maintains global consistency in dynamic environments where static, recorded-video calibration is impossible.

I. I NTRODUCTION Multi-camera tracking enables global identity consistency across camera networks, improving robustness under occlusion and expanding coverage. However, many multi-camera systems assume accurate 3D calibration (intrinsics/extrinsics, surveyed landmarks, or specialized procedures). In real-world deployments, relying on such precision is not merely expensive; it is often operationally infeasible [5]. Three primary barriers prevent the adoption of traditional calibration in large-scale facilities. First, privacy regulations (e.g., GDPR) often strictly prohibit the recording and storage of video footage required for offline calibration bundles. Second, bandwidth constraints in large retail or industrial networks preclude the transmission and synchronization of high-resolution video streams from hundreds of cameras to a central server. Third, covering thousands of square meters with physical calibration targets (e.g., checkerboards) is logistically impossible without disrupting site operations. We target deployments where operators can provide only manual homographies from image coordinates to a site map. Such homographies are fast to obtain but are typically imperfect due to occlusions, imperfect floor plans, and extrapolation

errors. Crucially, static homographies fail when cameras are inevitably nudged or moved during daily operations. Key idea. We propose a system that requires neither video recording nor physical site access. By refining initial manual homographies using a projection optimization (PO) procedure that operates on live detection metadata, we align multi-camera projections in BEV space ”on-the-fly.” This approach ensures privacy compliance and allows the system to autonomously heal geometric alignment errors caused by camera movement. a) Contributions: A zero-overhead projection optimization method that refines manual homographies asynchronously. By running exclusively on lightweight, clustered cross-view detection metadata, it continuously adapts to environmental changes without adding computational latency to the real-time pipeline. • A global BEV multi-camera tracker that associates multi-view observations using combined motion and appearance cues before fusing them into a shared geometric state, operating purely on privacy-safe metadata. • A trajectory-level cross-camera unification stage that merges duplicate global tracks using temporal alignment, motion dynamics, and appearance costs.

•

b) Scope of Tracking Framework: It is important to note that while we present an end-to-end tracking pipeline, our core scientific contribution is the on-the-fly Projection Optimization (PO) calibration module, not a novel tracking architecture. The custom BEV tracker described herein is provided strictly as an evaluation vehicle to demonstrate the tangible improvements in trajectory alignment resulting from our geometric refinement. Because PO operates purely on the coordinate mapping layer, any multi-camera tracker that relies on spatial gating or trajectory alignment will inherently benefit from the restored geometric fidelity our module provides. II. R ELATED W ORK A well-structured multi-camera tracking system must balance geometric precision with deployment feasibility. Below, we contextualize our approach against the three primary calibration paradigms in the tracking literature, highlighting their operational shortcomings and how our proposed system explicitly addresses them.

a) Multi-Camera Tracking with 3D Calibration: PreviIII. S YSTEM OVERVIEW ous work often assumes an accurate 3D camera calibration At each time step, each camera produces detections (boundto triangulate targets or fuse observations in a shared 3D ing boxes, confidence scores, and appearance embeddings). space [10], [11]. Obtaining such precision in the wild typically A chosen image point (e.g. leg point / bottom-center) is requires costly manual setup or specialized procedures, such mapped to BEV via a per-camera projection. Crucially, the edge as surveying architectural landmarks or placing physical nodes transmit only this lightweight metadata to the central checkerboards across the site [5]. Failure Mode: These systems tracker, not video frames. This design respects strict privacy are inherently brittle in active surveillance environments. If policies regarding video persistence and minimizes bandwidth a camera is nudged or moved during daily operations, the usage. We then (i) match per-camera detections to shared BEV rigorous static calibration is immediately invalidated, requiring tracks, (ii) fuse the matched spatial detections and appearance a costly site visit to repeat the setup. Our Solution: Our embeddings to continuously update the global track states, and system abandons the need for static, surveyed 3D calibration. (iii) periodically merge duplicate BEV tracks via trajectory By continuously running our Projection Optimization (PO) unification. on live metadata, the system autonomously heals geometric IV. M ETHOD alignment errors on-the-fly, gracefully recovering from camera A. Projection from Manual Homographies movements without human intervention. For camera i ∈ {1, . . . , N }, we initialize a projective map(0) ping (homography) Ai from the image coordinates to a shared b) Planar (Homography) Multi-Camera Tracking: To 2 bypass full 3D calibration, projecting multi-view observations BEV plane [5]. Given an image point p ∈ R , we compute 2 to a common ground plane (bird’s-eye view) using planar its projection as x = Π(p; Ai ) ∈ R , where Π denotes the followed by dehomogenization. homographies is a standard approach for fusing data [6], homogeneous projection (0) In practice, Ai is obtained from a small set of manual [4]. Failure Mode: These systems typically assume the provided homographies are highly precise. In practice, manual correspondences. While fast to annotate, these suffer from homographies frequently exhibit residual misalignment due well-known critical limitations: (i) down-scaling reduces point to extrapolation errors, lens distortion, and imperfect floor placement accuracy, (ii) occlusions force landmark estimation, plans [5]. This residual spatial scatter is catastrophic for (iii) imperfect floor plans mismatch physical reality, (iv) floortracking: it causes multi-view observations of the same target contact ambiguity arises due to perspective differences, (v) to fail spatial distance gating, severely fragmenting trajectories extrapolation errors grow significantly outside the annotated and causing rampant identity switches. Our Solution: We region, and (vi) lens distortion cannot be perfectly modeled treat manual homographies merely as a coarse warm start. By by simple projective transforms. In a static system, these errors dynamically clustering cross-view Re-ID correspondences, our are permanent; in our system, they serve merely as a warm centroid-anchored optimizer continuously pulls the projection start for online optimization. matrices into alignment. This drastically reduces spatial scatter B. Matching Point Acquisition and restores the integrity of the tracking system’s spatial gating The projection optimization described in Section IV-C relies logic. on clusters of corresponding points {Cc } observed across multiple views. To acquire these correspondences in live c) Weak Calibration and Self-Calibration: Methods for deployments where video recording is prohibited or infeasible, learning cross-view alignment or self-calibration from data we employ an online matching method driven by visual Retypically attempt to deduce geometry by observing moving Identification (Re-ID). We run an appearance-based tracker targets over time [3]. Failure Mode: These approaches gener- across the camera network to associate detections based on ally require strong environmental assumptions (e.g., vehicles embedding similarity. moving in perfectly straight lines to find vanishing points), Crucially, because projection optimization is highly sensitive centralized video processing, and long sequences of recorded to outliers, we prioritize precision over recall. We apply a footage to converge. While effective for traffic surveillance, strictly low matching distance threshold (or high similarity these assumptions fail completely in retail or pedestrian envi- score), rejecting any ambiguous associations. This ensures that ronments where targets move erratically. Furthermore, in large- we harvest a smaller set of highly robust correspondences scale deployments, strict privacy regulations (e.g., GDPR) and rather than a large volume of noisy matches, preventing false bandwidth constraints prohibit the recording and transmission associations from corrupting the spatial alignment. of the raw video required for these offline calibration algorithms. This online approach is critical for ”on-the-fly” adaptation. Our Solution: Our approach differs by being fully online If a camera is accidentally nudged, the Re-ID stream continues and privacy-compliant. It relies exclusively on lightweight, to find correspondences between the (now shifted) view and anonymized detection metadata to refine the geometry via a its neighbors. The subsequent optimization step naturally pulls stable BEV consistency objective. This eliminates the need the projection matrix back into alignment without requiring a for video transmission entirely, making it ideal for bandwidth- site visit or manual recalibration, ensuring long-term system constrained, privacy-sensitive networks. resilience.

a) Spatial Subsampling: Since the transformation be- An optimizer such as Adam is applied to the current camera tween the image plane and the ground plane is modeled as matrix parameters A(k) to minimize this loss L(k) across all i a homography (a smooth, planar mapping), dense clusters of clusters and calibration points. The centroid-based loss naturally points within a small image neighborhood provide redundant extends from two cameras to N cameras: any number of views geometric constraints. To improve computational efficiency that contribute points to a cluster are all pulled toward the same and prevent overfitting to local clusters, we apply spatial cluster centroid in BEV space, yielding the updated matrices subsampling on the image plane. We enforce a configurable A(k+1) for the next iteration. i minimum pixel distance between selected points. This ensures c) Slowly updated centroids: To avoid trivial collapse a uniform distribution of constraints across the field of view and excessive drift, the reference centroids µc(kref ) used in (k ) without unnecessary redundancy. Equation 2 are refreshed only intermittently. Initially, µc ref = (0) µc . Every K optimization iterations we recompute: C. Projection Optimization (PO) 1 X (k) (k ) µc ref,new = xp . (3) Nc p∈Cc

D. Multi-Camera Tracking in BEV

Fig. 1. Projection inconsistency before optimization. A single person is observed simultaneously across multiple overlapping camera views (left). When mapped into the shared bird’s-eye view (bottom right) using initial manual homographies, the individual camera projections (colored circles) exhibit significant spatial scatter around their computed centroid (black dot). The black arrow illustrates the multi-view alignment error that our projection optimization aims to minimize per camera.

a) Centroid-based optimization objective: As shown in Figure 1, initial manual homographies often result in significant spatial disagreement when the same object is viewed from different perspectives. To resolve this, calibration points are grouped into clusters {Cc }, where each cluster corresponds to one physical landmark observed in two or more cameras. Using (0) the original manual projection matrices Ai , we compute an initial BEV centroid for each cluster: 1 X (0) µ(0) xp , (1) c = Nc p∈Cc

where Nc = |Cc |, p represents an individual image-space observation (e.g., a bounding box foot point from a specific (0) camera) belonging to cluster Cc , and xp is its corresponding BEV projection under the original matrices. b) Optimization over camera matrices: Because the system continuously adapts on-the-fly, the projection matrix for each camera is updated iteratively. The optimization cost at iteration k is defined as the sum of the squared distances from each projected point to the reference centroid of its cluster: XX 2 (kref ) L(k) = x(k) . (2) p − µc c

p∈Cc

a) Design principle: a shared global track pool in BEV: The multi-camera (MC) tracker maintains a single set of global tracks whose latent state lives in a shared bird’s-eye-view (BEV) plane. Each camera acts as an independent measurement source that can update any global track. b) Per-camera inputs: At each timestamp t, each camera (i) (i) Mi i produces a set of detections Dt = {dt,m }m=1 with: (i) a 2D (i) image box and confidence score, (ii) a chosen image point pt,m that approximates ground contact (e.g. “leg point” / bottom(i) center), and (iii) an appearance embedding et,m . The image point is projected to BEV using the per-camera projection model (Section IV-A–IV-C), which yields BEV measurements (i) zt,m ∈ R2 . c) Refining Leg Points via Single-View Tracking: Since our system relies on planar homographies rather than full 3D calibration, multi-camera association is highly sensitive to input noise; a few pixels of vertical jitter in the image plane can translate into large spatial displacements in the BEV. We mitigate this by feeding the multi-camera tracker with smoothed state estimates from the single-view tracker rather than raw detections. This temporal stabilization reduces projection variance and provides better-calibrated uncertainty covariances, significantly improving the recall of cross-camera matching under tight Mahalanobis gates. 1) Per-camera association against the shared pool: a) Parallel matching: For each timestamp t, we first (j) (j) predict all global tracks to obtain {xt|t−1 , Pt|t−1 }j . For each camera i, we then independently match its BEV measurements (i) {zt,m }m to the same global track pool. This step can be parallelized across cameras, allowing for efficient multi-sensor fusion. b) Motion cost (Mahalanobis): For detection m at time t in camera i and track j, define innovation and covariance: (i)

(j)

(i,j) fm = zt,m − Hxt|t−1 ,

(j)

S(j) = HPt|t−1 H⊤ + R(j) , (4)

where R(j) may be adapted over time (below). The motion distance is (i,j) ⊤ (i,j) dmot (m, j) = (fm ) (S(j) )−1 fm .

(5)

c) Appearance cost: We use Euclidean distance between F. Implementation Details (i) the detection embedding et,m and the track embedding e(j) . Embeddings are computed following standard Re-ID praca) Initial Homography Estimation: The initial projection tices [9]. (0) matrices Ai are established via a dedicated annotation d) Combined cost and assignment: Since MC tracks a BEV point (not a 2D box), we do not use IoU-based interface that displays the camera view alongside a 2D site map. fallback matching. We construct a combined cost c(m, j) = Operators select a minimum of four corresponding point pairs α dmot (m, j) + β dapp (m, j). After applying gating thresholds (anchors) between the image plane and the bird’s-eye-view for both motion and total cost, we resolve the one-to-one (BEV) map, deliberately distributing these points as widely as possible across the field of view to maximize coverage and assignment for camera i using the Hungarian algorithm [7]. minimize extrapolation errors. The exact number of annotated 2) Cross-camera fusion into the global state: After parallel points is strictly scene-dependent, varying based on the matching, a global track j may receive concurrent measureavailability of distinct physical landmarks. To ensure a robust ments from multiple cameras. We fuse these observations initialization, the tool provides a real-time interactive preview: via sequential Kalman measurement updates [2], where the as the operator hovers over the image, the corresponding BEV posterior state and covariance from one camera’s update projection is actively displayed. This feedback loop allows immediately serve as the prior for the next. To handle users to iteratively add, move, or delete anchor points until a heterogeneous camera error (e.g., varying projection variance satisfactory baseline alignment is achieved. or detector noise), we adapt the measurement noise covariance (j) R online based on recent residual magnitudes. Finally, the b) Re-ID Model and Matching Policy: For cross-camera global track’s appearance embedding e(j) is updated via an visual association in the live deployment (Section IV-B), the exponential moving average (EMA) over the matched multi- system utilizes a PLR-OSNet [12] backbone to extract 2560view detections. dimensional appearance embeddings. Given a frame processing 3) Track initiation and lifecycle: rate of 5 fps, object detections from different cameras are tema) Unmatched detections and new track creation: After porally synchronized by matching their metadata timestamps association and fusion, each camera may have unmatched within a narrow tolerance window of ±100 ms. To prioritize detections. We reuse the same logic as single-view for new precision over recall during point acquisition, cross-camera track initiation: only detections above a class-specific opening associations are evaluated using Euclidean distance and require threshold can spawn new tracks. This prevents low-confidence a strict similarity threshold. clutter from creating global IDs. Tracks transition from tentative c) Outlier Rejection: To further mitigate false Re-ID to confirmed on the basis of configured hit/miss logic. matches and prevent erroneous associations from corrupting the geometric alignment, we enforce a strict, configurable spatial E. Trajectory-Level Cross-Camera Track Unification gating mechanism. Any projected point that yields a spatial distance greater than a defined threshold from its calculated To resolve temporary track fragmentation (e.g., during crosscluster centroid in the shared bird’s-eye-view space is rejected camera handovers or severe occlusions), we run a trajectoryas an outlier and excluded from the optimization constraints. level merge stage after the per-frame fusion update. We first filter candidates using a strict physical gate: tracks observed by d) Optimization Parameters and Scheduling: The projecthe same camera at the same timestamp are assigned infinite tion optimization is performed using the Adam optimizer. Becost, preventing the collapse of distinct neighboring objects cause the BEV projection is highly sensitive to small perturbainto a single identity. tions in the homography matrix—where minor parameter shifts For admissible candidate pairs (i, j) with a set of overlapping can cause large spatial displacements and lead to divergence— timestamps Ωij , we aggregate their historical trajectories to we employ a very low learning rate of 1 × 10−5 . We run assess similarity. Rather than relying on a single frame, we the optimization for 300,000 epochs; because it optimizes compute the average relative state difference d̄ij and the only a minimal parameter space (a 3 × 3 matrix per camera) average combined covariance Σ̄ij across the entire temporal over lightweight metadata, this executes rapidly. To stabilize overlap. The structural motion cost is then evaluated via the convergence, the reference cluster centroids (Equation 3) are −1 Mahalanobis distance dmotion = d̄⊤ updated every 15,000 epochs, but these updates are intentionally ij ij Σ̄ij d̄ij . This motion cost is fused with a cosine distance demb frozen after the first 30,000 epochs (10% of the schedule) to ij between the tracks’ appearance embeddings. Pairs exceeding allow the matrix parameters to settle into a stable geometric configured spatial or appearance thresholds are discarded, and state. Furthermore, to avoid any computational bottleneck, this the remaining valid pairs are resolved via greedy assignment. optimization service runs asynchronously (e.g., once per day) Upon merging, the retained track absorbs the duplicate, on a separate machine using the already-transmitted metadata. updating its appearance embedding via a weighted average This architecture ensures that continuous geometric refinement proportional to the number of multi-view detections each track adds zero computational overhead or latency to the real-time has accumulated. multi-view tracking pipeline.

V. E XPERIMENTS

D. Results

a) Quantitative Analysis and the Impact of Alignment: Quantitative results for Datasets A, B, and C are presented We evaluated on three proprietary multi-camera datasets in Table I. Across all environments, replacing raw manual collected in real retail environments. Due to privacy and homographies with our Projection Optimization (PO) yields business constraints, the raw videos and annotations cannot consistent improvements in both geometric fidelity and tracking be publicly released. All datasets are recorded at 5 FPS with consistency. fixed cameras and are time-synchronized. The ground truth The direct correlation between BEV alignment error and is labeled with global identities across all cameras: the same tracking stability is profoundly demonstrated when comparing person observed in multiple views is assigned the same GT the dynamics of Datasets B and C. Prior to optimization, track ID. both datasets suffered from a substantial BEV alignment a) Dataset A (Supermarket): This dataset contains 11 error of 42 cm. However, the manifestation of this error cameras and spans 2:20 minutes. Cameras are installed at depends heavily on scene density. Dataset B represents a busy approximately 45◦ view angle and around 3 m height. It mall environment with high pedestrian density (80 unique includes 53 global ground-truth tracks. trajectories). In crowded scenes, a 42 cm projection discrepancy b) Dataset B (Mall): This dataset contains 5 cameras and is catastrophic: it not only pushes cross-camera observations spans 3:50 minutes. Cameras are installed at approximately outside their correct Mahalanobis matching gate (Equation 5), 40◦ view angle and around 4.5 m height. It includes 80 global but frequently pushes them into the spatial gates of neighboring ground-truth tracks. pedestrians. This leads to rampant identity mixing, resulting in c) Dataset C (General store): This dataset contains 8 an extremely high identity switch count (182 IDSW) under the cameras and spans 2:50 minutes. Cameras are installed at manual baseline. By applying our centroid-anchored PO, the approximately 45◦ view angle and around 4 m height. It alignment error in Dataset B is reduced threefold to 14 cm. This includes 45 global ground-truth tracks. spatial convergence resolves the identity collisions, directly driving the dramatic reduction in ID switches down to 38, B. Evaluation Metrics while boosting IDF1 from 0.717 to 0.892. Interestingly, Dataset C presents a different but equally We report multi-object tracking metrics (e.g., IDF1 / HOTA / MOTA) [10], [8], [1], ID switches, and calibration/alignment important insight. Despite starting with the same 42 cm error, its lower target density (a general store with 45 trajectories) metrics. means that spatial offsets rarely caused projections to collide • BEV alignment error: for each landmark cluster, compute its BEV centroid as the mean of all projected points in that with neighboring identities. Consequently, the raw ID switch cluster; report the mean ℓ2 distance from each projected count remained relatively stable (25 to 24). However, the 42 cm point to its cluster centroid, averaged over all points across error still caused the tracker to drop associations and break trajectories. By reducing the alignment error from 42 cm to all clusters. 24 cm, PO ensures that targets maintain their correct global • Tracking: IDF1, HOTA; ID switches. identity for much longer durations without fragmenting, which is reflected in the significant improvements in both IDF1 (0.806 C. Baselines and Evaluation Strategy to 0.906) and HOTA (0.683 to 0.810). This confirms that Our experimental objective is to isolate and quantify the mitigating projection scatter is the primary mechanism for impact of geometric calibration on multi-camera association. maintaining global trajectory purity, regardless of scene density. Therefore, rather than comparing disparate tracking architectures—which conflates calibration quality with variations in TABLE I data association heuristics—we evaluate the identical multiQ UANTITATIVE T RACKING AND A LIGNMENT R ESULTS . camera tracking pipeline under two different geometric conditions: Dataset Method IDF1 ↑ HOTA ↑ IDSW ↓ AlignErr ↓ • Manual homographies (no PO). We use the operatorManual 0.938 0.864 7 37 cm provided per-camera homographies as-is to project detec- A + PO (ours) 0.966 0.926 0 20 cm tions into BEV and run the multi-camera tracker. This Manual 0.717 0.544 182 42 cm baseline measures how much tracking performance is lost B + PO (ours) 0.892 0.783 38 14 cm when a standard system relies solely on raw, unrefined manual calibration. Manual 0.806 0.683 25 42 cm C • Manual + PO (ours). Starting from the exact same manual + PO (ours) 0.906 0.810 24 24 cm homographies, we run projection optimization to refine the per-camera projections, and then run the identical multiVI. D ISCUSSION AND L IMITATIONS camera tracker. By keeping the tracking logic completely static, any improvement over the previous baseline is a) Planar Ground Assumption: Our homography-based strictly attributable to better cross-camera BEV alignment. formulation strictly assumes that all targets move on a dominant A. Setup and Datasets

(a) Before Optimization

(b) After Optimization

Fig. 2. Detailed anatomy of a single spatial cluster before (a) and after (b) Projection Optimization. In this instance, the blue point represents a target’s projection from Camera A, while the yellow point represents the exact same target at the same timestamp from Camera B. The central black point designates the computed cluster centroid, with the grey lines visualizing the spatial distance from each camera’s projection to this shared center. The number “2” denotes that this specific cluster fuses observations from two overlapping cameras. Crucially, while this highlights a single synchronized instance, the full-scale maps (Figure 3) plot all clusters generated across the entire video timeline for all people. Because the optimization is agnostic to time and global identity—focusing purely on pulling corresponding multiview points toward their local centroids—these macroscopic maps effectively visualize the complete geometric input space and convergence objective of the PO algorithm.

(a) Before Optimization

(b) After Optimization

c) Initialization Sensitivity: Finally, centroid-based optimization is designed as a local refinement step rather than a global geometric solver. It assumes that the initial manual homographies provide a roughly correct starting point regarding scale and orientation. If the initial manual input is grossly incorrect (e.g., flipped axes or massive scale errors), the optimization may converge to a degenerate local minimum. However, in practice, the real-time visual feedback provided by our annotation UI (Section IV-F) effectively prevents these severe initialization errors, ensuring the starting geometry is well within the convergence basin of the PO algorithm. VII. C ONCLUSION We presented a calibration-light multi-camera tracking framework that successfully bridges the gap between geometric precision and real-world deployment constraints. By introducing a centroid-anchored Projection Optimization (PO) method, we demonstrated that coarse, manual planar homographies can be dynamically refined into a robust spatial coordinate system using exclusively live, anonymous detection metadata. This continuous, on-the-fly alignment eliminates multi-view projection scatter, restoring the integrity of spatial distance gating and significantly reducing global identity switches. Furthermore, because the optimization operates asynchronously on lightweight metadata, it completely decouples geometric self-healing from the real-time tracking loop, incurring zero computational overhead. Ultimately, our approach yields a highly scalable, privacy-compliant pipeline capable of maintaining stable global trajectories in dynamic environments where traditional static 3D calibration and centralized video processing are fundamentally prohibitive or infeasible.

Fig. 3. Global BEV projection consistency across the entire video timeline before (a) and after (b) optimization.

R EFERENCES

ground plane (z = 0). Although this approximation holds for the majority of retail and indoor environments, the projection model is inherently limited in scenes with significant elevation changes, such as ramps or split-level flooring. In such nonplanar scenarios, a single homography cannot accurately map foot points to the bird’s-eye view. Future work could address this by modeling complex sites as piecewise-planar spaces, or by replacing the rigid planar matrix with a lightweight neural projection model capable of learning non-linear, terrain-aware mappings directly from the sparse detection metadata. b) Dependence on View Overlap: The efficacy of Projection Optimization (PO) relies on the existence of a connected graph of overlapping views. The optimization constraints are derived entirely from joint observations—either shared static landmarks or clustered detections visible in multiple cameras. If a camera is spatially isolated (i.e., it shares no common field of view with the network), its projection parameters cannot be geometrically refined relative to the global frame. In these disjoint cases, the system safely falls back to the initial manual calibration, relying more on the visual Re-ID embeddings rather than spatial gating for cross-zone associations.

[1] Keni Bernardin and Rainer Stiefelhagen. Evaluating multiple object tracking performance: the clear mot metrics. EURASIP Journal on Image and Video Processing, 2008(1):246309, 2008. [2] Robert Grover Brown and Patrick YC Hwang. Introduction to random signals and applied kalman filtering: with matlab exercises and solutions. Introduction to random signals and applied Kalman filtering: with MATLAB exercises and solutions, 1997. [3] Markéta Dubská, Adam Herout, Roman Juránek, and Jakub Sochor. Fully automatic roadside camera calibration for traffic surveillance. IEEE Transactions on Intelligent Transportation Systems, 16(3):1162–1171, 2014. [4] Francois Fleuret, Jerome Berclaz, Richard Lengagne, and Pascal Fua. Multicamera people tracking with a probabilistic occupancy map. IEEE transactions on pattern analysis and machine intelligence, 30(2):267–282, 2008. [5] Richard Hartley and Andrew Zisserman. Multiple view geometry in computer vision. Cambridge university press, 2003. [6] Saad M Khan and Mubarak Shah. Tracking multiple occluding people by localizing on multiple scene planes. IEEE transactions on pattern analysis and machine intelligence, 31(3):505–519, 2008. [7] Harold W Kuhn. The hungarian method for the assignment problem. Naval research logistics quarterly, 2(1-2):83–97, 1955. [8] Jonathon Luiten, Aljosa Osep, Patrick Dendorfer, Philip Torr, Andreas Geiger, Laura Leal-Taixé, and Bastian Leibe. Hota: A higher order metric for evaluating multi-object tracking. International journal of computer vision, 129(2):548–578, 2021. [9] Hao Luo, Youzhi Gu, Xingyu Liao, Shenqi Lai, and Wei Jiang. Bag of tricks and a strong baseline for deep person re-identification. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, 2019.

[10] Ergys Ristani, Francesco Solera, Roger Zou, Rita Cucchiara, and Carlo Tomasi. Performance measures and a data set for multi-target, multicamera tracking. In European conference on computer vision, pages 17–35. Springer, 2016. [11] Zheng Tang, Milind Naphade, Ming-Yu Liu, Xiaodong Yang, Stan Birchfield, Shuo Wang, Ratnesh Kumar, David Anastasiu, and Jenq-Neng Hwang. Cityflow: A city-scale benchmark for multi-target multi-camera vehicle tracking and re-identification. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8797–8806, 2019. [12] Ben Xie, Xiaofu Wu, Suofei Zhang, Shiliang Zhao, and Ming Li. Learning diverse features with part-level resolution for person re-identification. In Chinese Conference on Pattern Recognition and Computer Vision (PRCV), pages 16–28. Springer, 2020.

Record · ID 965468 · SHA-256 abe5dda03aa98e71
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.