# Traffic3D: Edge-Aware Primitive Decomposition with Graph Relational Refinement for Lightweight Monocular 3D Traffic Scene Reconstruction --- **Authors:** Research Proposal **Keywords:** Monocular 3D Reconstruction, Semantic Segmentation, Scene Graph Generation, Graph Neural Networks, Geometric Primitives, Point Cloud Generation, Autonomous Driving --- ## Abstract We present **Traffic3D**, a modular and lightweight pipeline for reconstructing semantically consistent 3D traffic scenes from a single monocular RGB image. Unlike existing methods that rely on dense voxel grids, multi-sensor fusion, or computationally expensive transformers, Traffic3D employs a five-stage cascade combining (i) edge-aware input augmentation with positional depth priors, (ii) a lightweight UNet with boundary-conditioned supervision for crisp semantic segmentation, (iii) PCA-based geometric primitive extraction with automatic scene graph construction, (iv) a parameter-efficient Graph Neural Network (<500K parameters) for relational refinement of scene elements, and (v) surface-sampled 3D point cloud generation. Our design philosophy prioritizes *real-time feasibility* (≥15 FPS on an RTX 3090) and *monocular-only operation*, eliminating dependencies on LiDAR, stereo cameras, or depth sensors. We provide a complete mathematical formulation of each stage, a novel edge-weighted loss function combining cross-entropy with on-the-fly boundary supervision, and three GNN variants (EdgeAwareSAGE, GATv2, and a learned Hybrid) for ablation. Preliminary benchmarks on synthetic traffic scenes demonstrate that all GNN variants satisfy the 500K parameter constraint (28.7K–62.0K), the full pipeline totals 4.36M parameters, and the system produces semantically labeled 3D point clouds with 2K–20K points per scene. We target evaluation on Cityscapes, BDD100K, and CARLA, with expected improvements of ≥15% in boundary IoU over non-edge-aware baselines and competitive 3D IoU (~0.68), centroid accuracy (~0.49m), and Chamfer distance (~0.041). --- ## 1. Introduction ### 1.1 Motivation Understanding the three-dimensional structure of traffic scenes is a cornerstone capability for autonomous driving, urban planning, augmented reality navigation, and robotic delivery systems. While LiDAR-based methods achieve high geometric fidelity [1, 2], they require expensive sensor hardware and produce sparse point clouds that lack semantic context. Multi-camera systems [3, 4] reduce cost but still impose calibration complexity and physical mounting constraints. The challenge of reconstructing dense, semantically labeled 3D scenes from a *single* RGB image—the most ubiquitous and inexpensive sensor modality—remains an open and compelling research problem. Recent advances in monocular 3D semantic scene completion (SSC) have demonstrated that single-image 3D reconstruction is feasible. MonoScene [5] pioneered the projection of 2D features into 3D voxel space using successive UNets, achieving 11.08 mIoU on SemanticKITTI. VoxFormer [6] improved this to 13.35 mIoU using sparse voxel queries and masked autoencoder-style densification. However, these methods operate on dense 256×256×32 voxel grids, requiring substantial computation (VoxFormer: ~4.4 FPS) and memory that prohibits real-time deployment. Concurrently, the scene graph generation (SGG) community has shown that explicit relational modeling between objects dramatically improves scene understanding [7, 8, 9]. Yet SGG methods typically rely on pre-computed object detections and operate in 2D, missing the rich 3D spatial relationships (occlusion ordering, support relations, proximity) that govern traffic scenes. We identify three critical gaps in the current literature: 1. **Dense-to-sparse gap**: Existing 3D SSC methods produce dense voxel occupancy grids, wasting computation on empty space. Traffic scenes are inherently composed of discrete objects (vehicles, pedestrians, infrastructure) that admit compact primitive representations. 2. **Boundary degradation**: Standard segmentation networks produce blurry boundaries between adjacent classes [10, 11], directly corrupting downstream primitive extraction when class regions bleed into neighbors. 3. **Relational modeling gap**: Current monocular 3D methods reconstruct objects independently, ignoring spatial relationships (a car *on* a road, a pedestrian *near* a crosswalk) that constrain physically plausible scene configurations. ### 1.2 Contributions Traffic3D addresses all three gaps through the following contributions: 1. **Edge-aware segmentation with boundary supervision**: We introduce a combined loss function `L = L_CE^edge + λ·L_boundary` that applies spatially varying cross-entropy weights using Sobel/Canny edge confidence maps and adds auxiliary boundary supervision derived on-the-fly from segmentation labels (no extra annotations required). Following the SBCB framework [10], the boundary head is discarded at inference, adding zero computational overhead. 2. **Primitive-based 3D scene abstraction**: Instead of dense voxels, we decompose the scene into semantically typed geometric primitives—cuboids for vehicles and buildings (PCA-fitted), cylinders for pedestrians, cones for vegetation, and planes for road and sky—providing an orders-of-magnitude more compact representation. 3. **Lightweight GNN relational refinement**: We design three GNN architectures (EdgeAwareSAGE, GATv2, Hybrid) all under 500K parameters that operate on the constructed scene graph to refine primitive features through message passing, encoding spatial relationships that improve scene consistency. 4. **Complete modular pipeline**: We deliver an end-to-end system from RGB input to colored, labeled 3D point cloud output with PLY export capability, a 4-phase training strategy, comprehensive evaluation metrics, and an ablation study framework—all in a single, documented codebase. ### 1.3 Design Requirements Our pipeline is designed under the following constraints, reflecting deployment realities: | Requirement | Target | Rationale | |---|---|---| | Input modality | Single RGB image | Cost, ubiquity | | Inference speed | ≥15 FPS (RTX 3090) | Real-time driving | | GNN parameters | <500K | Edge device feasibility | | Total parameters | <10M | Mobile deployment | | Point cloud density | 2K–20K points | Downstream compatibility | | Annotation dependency | Segmentation only | Reduce labeling cost | --- ## 2. Related Work ### 2.1 Monocular 3D Semantic Scene Completion The task of inferring dense 3D geometry and semantics from 2D images was formalized by Song et al. [12] using depth maps. **MonoScene** [5] (CVPR 2022) first demonstrated SSC from a single RGB image by projecting 2D EfficientNet-B7 features into 3D space via Feature Line-of-Sight Projection (FLoSP) and processing them with a 3D UNet. It achieved 34.16 IoU (geometry) and 11.08 mIoU (semantics) on SemanticKITTI, using a composite loss of cross-entropy, scene-class affinity, multi-scale geometric/semantic losses, and frustum proportion loss. **VoxFormer** [6] (CVPR 2023) improved upon MonoScene by adopting a two-stage sparse-to-dense strategy: (i) depth estimation produces a sparse set of visible occupied voxels as queries, then (ii) a deformable attention transformer with masked autoencoder-style self-attention densifies the scene. This achieved 44.02 IoU and 13.35 mIoU on SemanticKITTI, a +19.6%/+18.1% improvement, but at 2×10⁴ voxel queries it requires significant computation. More recently, **OccNet** [13] and **SurroundOcc** [14] have pushed multi-camera occupancy prediction to 39+ mIoU, while **SuperOcc** [15] (2025) introduced superquadric primitives as an alternative to dense voxels, achieving 36.1 mIoU at 30.3 FPS with 600 superquadric queries on the Occ3D benchmark—demonstrating that primitive-based representations can match dense methods while being significantly faster. Traffic3D draws inspiration from SuperOcc's primitive philosophy but operates in the monocular single-camera regime and uses explicit geometric primitive fitting rather than learned superquadric parameters, enabling interpretable scene abstractions without requiring occupancy supervision. ### 2.2 Real-Time Semantic Segmentation Real-time segmentation is essential for our pipeline since it constitutes the computational bottleneck. **BiSeNet** [16] introduced the two-stream architecture separating spatial and semantic paths. **STDC-Seg** [11] (CVPR 2021) demonstrated that this two-stream overhead is unnecessary: by designing the Short-Term Dense Concatenate (STDC) module with geometric channel decay, it achieved 76.8 mIoU at 97.0 FPS on Cityscapes (1080Ti, 768×1536). Critically, STDC-Seg uses a Detail Aggregation module with binary boundary supervision at 1/8 scale during training, showing that boundary-aware training improves segmentation quality without inference overhead. For lightweight architectures, the classic **UNet** [17] remains highly competitive when channel widths are reduced. With a base channel width of 32, a 4-level UNet achieves ~4.3M parameters while maintaining full-resolution output with skip connections—making it an ideal backbone for our pipeline where the 5-channel augmented input (RGB + positional + edge) requires a custom input layer. ### 2.3 Boundary-Aware Segmentation Standard cross-entropy loss treats all pixels equally, causing boundary degradation between adjacent classes. Several approaches address this: **SegFix** [18] (ECCV 2020) proposed a post-processing boundary refinement module that replaces predictions on boundary pixels with those from nearby interior pixels, improving boundary F-score by ~4 points on Cityscapes. **InverseForm** [19] showed that learning the inverse distance transform at boundaries improves boundary quality. Most relevant to our work, **SBCB** (Semantic Boundary Conditioned Backbone) [10] (2023) demonstrated that adding a lightweight Semantic Boundary Detection (SBD) head (e.g., CASENet, DFF, or DDS) during training—discarded at inference—improves segmentation mIoU from 79.5 to 80.3 on Cityscapes with DeepLabV3+ and from 80.1 to 82.0 with HRNet-48, achieving boundary F-scores of 74.4–78.9. Crucially, boundary ground truth is generated *on-the-fly* from segmentation labels, requiring no additional annotations. **BPKD** [20] (2023) similarly showed that boundary-privileged knowledge distillation, separating body and edge supervision, achieves state-of-the-art distillation results across architectures. Traffic3D adopts the SBCB principle: we add an auxiliary boundary prediction head that uses on-the-fly boundary GT from label transitions, combined with a novel *edge-weighted* cross-entropy where the Sobel/Canny edge confidence map spatially modulates the loss, focusing learning capacity on boundary regions. ### 2.4 Scene Graph Generation Scene Graph Generation (SGG) constructs a structured graph representation where nodes are objects and edges encode their pairwise relationships. **Neural Motifs** [7] (CVPR 2018) showed that simple frequency baselines are surprisingly competitive due to strong dataset biases. **VCTree** [8] introduced tree-structured contexts. More recently, **REACT** [21] (2024) achieved real-time SGG (23ms latency) using a decoupled YOLOv8 backbone with efficient prototype learning, demonstrating that SGG can be made practical for driving applications. For driving-specific topology, **TopoNet** [22] (2023) introduced a scene graph neural network for road topology reasoning, showing that graph-based representations improve lane-traffic element assignment in the OpenLane-V2 benchmark. In the generative domain, **DiffuseSG** [23] (2024) demonstrated joint diffusion of scene graphs and images, while recent work on **controllable 3D outdoor scene generation** [24] (2025) showed that scene graphs can condition diffusion models for generating entire 3D urban environments. Traffic3D builds scene graphs from segmentation-derived primitives rather than from object detections, enabling construction without a separate detection model and ensuring consistency between the segmentation and graph representations. ### 2.5 Primitive-Based 3D Representations Representing 3D scenes as compositions of geometric primitives has a long history in computer vision. **Differentiable Blocks World** [25] (NeurIPS 2023) optimizes textured superquadric meshes via differentiable rendering, producing compact scene decompositions from single images. **Marching-Primitives** [26] (2023) extracts superquadrics directly from signed distance functions, achieving state-of-the-art abstraction on ShapeNet. **SuperDec** [27] (2025) extended this to full scene-level decomposition using superquadric primitives on ScanNet++ and Replica scenes. For occupancy prediction, **SuperOcc** [15] parameterizes each primitive as `S = (m, r, s, ε, σ, c)` encoding center, rotation, scale, squareness, opacity, and semantic logits, using 600 queries to achieve competitive occupancy prediction at 30+ FPS. Traffic3D uses simpler geometric primitives (cuboids, cylinders, cones, planes) rather than superquadrics, trading some representational flexibility for interpretability and computational simplicity. Our PCA-based fitting operates directly on segmentation masks without requiring learned primitive parameters or differentiable rendering. ### 2.6 Monocular Depth Estimation Depth estimation is fundamental for lifting 2D observations to 3D. **Depth Anything V2** [28] (2024) established a new foundation model for monocular depth estimation, training on synthetic data and achieving state-of-the-art results across benchmarks with models ranging from 24.8M (ViT-S) to 335.3M (ViT-L) parameters. **Depth Pro** [29] (2024) achieves metric depth with sharp boundaries in 0.3 seconds per 2.25MP image. **AnyDepth** [30] (2025) introduced a lightweight alternative using DINOv3 as a visual encoder with a Simple Depth Transformer decoder, achieving competitive accuracy with reduced computational overhead. Traffic3D uses a simplified vertical positional encoding as a depth prior rather than a learned depth estimator, trading metric accuracy for zero additional computation. This is motivated by the observation that in forward-facing driving cameras, vertical image position correlates strongly with depth (top=far, bottom=near). We identify integration of a lightweight metric depth model (e.g., Depth Anything V2-Small) as a key future extension. ### 2.7 Graph Neural Networks The GNN architectures in Traffic3D are grounded in two foundational methods: **GraphSAGE** [31] (NeurIPS 2017) introduced inductive node representation learning via fixed-size neighborhood sampling and aggregation, achieving 0.950 F1 on citation networks. The mean aggregator is fastest while remaining near-optimal. However, standard SAGEConv does not support edge features. **GATv2** [32] (ICLR 2022) proved that the original GAT [33] computes *static* attention (independent of the query node), while GATv2 computes *dynamic* attention via `a^T LeakyReLU(W[h_i||h_j])`, which is strictly more expressive. GATv2 natively supports edge features via the `edge_dim` parameter, making it ideal for scene graphs where edge attributes (distance, relative position) carry semantic meaning. Traffic3D introduces **EdgeAwareSAGEConv**, a custom MessagePassing layer that injects edge features into GraphSAGE via additive projection (`m_ij = W_n·h_j + W_e·e_ij`), following the FiLM-style injection pattern from EGNAS [34]. We additionally propose a **Hybrid GNN** that combines SAGE and GAT with a learned sigmoid gate, automatically blending neighborhood aggregation and attention-based message passing. --- ## 3. Methodology ### 3.1 Pipeline Overview Let `I ∈ ℝ^{H×W×3}` denote an input RGB image. Traffic3D processes `I` through five sequential stages: ``` I ──→ [Stage 1] ──→ F ──→ [Stage 2] ──→ S ──→ [Stage 3] ──→ G ──→ [Stage 4] ──→ G' ──→ [Stage 5] ──→ P Augment Segment Extract Refine Generate (H×W×5) (H×W×K) Graph Graph' PointCloud ``` where `F` is the augmented 5-channel tensor, `S` is the semantic segmentation, `G = (V, E)` is the scene graph, `G'` is the refined graph, and `P` is the output point cloud. ### 3.2 Stage 1: Input Augmentation We augment the RGB input with two additional channels encoding domain-specific inductive biases: **Positional Encoding.** For forward-facing driving cameras, vertical image position is a strong proxy for depth (the "ground plane prior"). We encode this as: ``` P(x, y) = y / H, where y ∈ [0, H-1] ``` yielding `P = 0` (top, far) to `P = 1` (bottom, near). This provides the subsequent UNet with an explicit depth cue without requiring a separate depth estimator. **Edge Confidence Map.** We compute a differentiable edge confidence map `C(x, y) ∈ [0, 1]` using either Sobel or Canny operators: *Sobel variant:* ``` C_sobel(x, y) = normalize(√(G_x² + G_y²)) ``` where `G_x = I_gray * K_sobel_x` and `G_y = I_gray * K_sobel_y` are the horizontal and vertical Sobel gradients of the luminance image `I_gray = 0.299R + 0.587G + 0.114B`. *Differentiable Canny variant:* ``` I_smooth = I_gray * G_σ (Gaussian blur, σ=1.0) M = √(G_x(I_smooth)² + G_y(I_smooth)²) (gradient magnitude) C_canny = 0.7·σ_s(τ(M - θ_high)) + 0.3·σ_s(τ(M - θ_low)) ``` where `σ_s` is the sigmoid function and `τ = 20` controls threshold sharpness. This provides a differentiable approximation of the non-maximum suppression and double thresholding steps of classical Canny. The augmented input is: ``` F(x, y) = [R(x,y), G(x,y), B(x,y), P(x,y), C(x,y)] ∈ ℝ^{H×W×5} ``` The edge confidence map `C` serves dual purposes: (i) as an input channel providing explicit boundary information to the segmentor, and (ii) as a weight map for the edge-weighted loss function (Section 3.7). ### 3.3 Stage 2: Edge-Weighted Semantic Segmentation **Architecture.** We employ a lightweight UNet with 4 encoder and 4 decoder stages: ``` Encoder: Decoder: ┌──────────────────┐ ┌──────────────────┐ │ DoubleConv(5→f) │ ─── skip ────→ │ Up(2f→f) + skip │ → Conv1×1(f→K) ├──────────────────┤ ├──────────────────┤ │ Down(f→2f) │ ─── skip ────→ │ Up(4f→2f) + skip │ ├──────────────────┤ ├──────────────────┤ │ Down(2f→4f) │ ─── skip ────→ │ Up(8f→4f) + skip │ ├──────────────────┤ ├──────────────────┤ │ Down(4f→8f) │ ─── skip ────→ │ Up(16f→8f) + skip│ ├──────────────────┤ └──────────────────┘ │ Down(8f→8f) │ ── bottleneck (Dropout2d) └──────────────────┘ ``` where `f = 32` (lightweight) or `f = 64` (standard), and each `DoubleConv` consists of `[Conv3×3 → BN → ReLU]×2`. With `f = 32` and bilinear upsampling, the UNet totals ~4.33M parameters. The classification head produces logits `S ∈ ℝ^{B×K×H×W}` via a 1×1 convolution. An auxiliary **boundary head** (training only) produces `B_pred ∈ [0,1]^{B×1×H×W}` via `[Conv3×3 → BN → ReLU → Conv1×1 → Sigmoid]`. **Edge Weighting.** We apply the edge confidence map to amplify segmentation logits at boundaries: ``` S'(x, y) = S(x, y) · (1 + α · C(x, y)) ``` where `α = 1.0` controls amplification strength. Non-edge regions (where `C ≈ 0`) are unaffected, while boundary regions receive amplified logits, encouraging more decisive class assignments at transitions. ### 3.4 Stage 3: Primitive Extraction Given the segmentation map `seg(x, y) = argmax_k S'(x, y, k)`, we extract geometric primitives through three steps: **Step 1: Connected Component Analysis.** For each class `k`, we compute the binary mask `M_k = (seg == k)` and extract connected components via flood-fill (using scipy.ndimage.label). Components below a minimum size threshold `τ_min = 100` pixels are discarded. **Step 2: Class-to-Primitive Mapping.** Each Cityscapes class is associated with a primitive type: | Class | Primitive | Fitting Method | |---|---|---| | car, truck, bus, building, wall, fence | Cuboid | PCA | | person, rider, pole | Cylinder | Extent-based | | vegetation | Cone | Extent-based | | road, sidewalk, sky, terrain | Plane | PCA normal | **Step 3: Primitive Fitting.** For each connected component, we project pixel coordinates to approximate 3D coordinates using the positional depth prior: ``` Z(y) = d_scale · (1 - y/H) + Z_min X(x, y) = (x - W/2) · Z(y) / f_approx Y(x, y) = (y - H/2) · Z(y) / f_approx ``` where `d_scale` is a class-dependent depth scale, `Z_min = 2.0` m, and `f_approx = W` (approximate focal length). **Cuboid fitting** applies PCA to the 3D point set: the eigenvectors define orientation (converted to quaternion), and the eigenvalues determine spatial extent (`size = 2√λ`). The rotation matrix is corrected to ensure `det(R) = +1`. **Cylinder/Cone fitting** computes axis-aligned extents from the 3D point range, with the cylinder radius as `max(Δx, Δz)/2` and height as `Δy`. **Plane fitting** uses the smallest PCA eigenvector as the surface normal, with extent determined by the spatial range. Each primitive `p_i` is characterized by: ``` p_i = (type_i, class_i, centroid_i ∈ ℝ³, size_i ∈ ℝ³, orientation_i ∈ ℝ⁴) ``` where `orientation` is a unit quaternion `[w, x, y, z]`. ### 3.5 Stage 3b: Scene Graph Construction We construct a graph `G = (V, E, X, A)` where: **Nodes** `V = {v_1, ..., v_N}` correspond to primitives. Each node feature is: ``` h_i = [emb(class_i) ∈ ℝ^16 ‖ centroid_i ∈ ℝ³ ‖ size_i ∈ ℝ³ ‖ orientation_i ∈ ℝ⁴] ∈ ℝ^26 ``` where `emb(·)` is a learned embedding layer mapping class indices to 16-dimensional vectors. **Edges** `E = {(i,j) : ‖centroid_i - centroid_j‖₂ < d_thresh}` connect nodes within a distance threshold `d_thresh = 5.0` m. If no edges are created (isolated node), we fall back to connecting each node to its nearest neighbor. **Edge features** encode spatial relationships: ``` e_ij = [dist_ij ∈ ℝ ‖ adj_ij ∈ {0,1} ‖ Δpos_ij ∈ ℝ³] ∈ ℝ⁵ ``` where `dist_ij = ‖c_i - c_j‖₂`, `adj_ij = 𝟙(dist_ij < d_thresh/2)`, and `Δpos_ij = c_j - c_i`. ### 3.6 Stage 4: GNN Relational Refinement We propose three GNN architectures, all satisfying the <500K parameter constraint: #### 3.6.1 EdgeAwareSAGEConv Standard SAGEConv [31] does not accept edge attributes. We define a custom MessagePassing layer: ``` Message: m_{j→i} = W_n · h_j + W_e · e_{ij} Aggregate: ā_i = MEAN({m_{j→i} : j ∈ N(i)}) Update: h_i^{(l+1)} = W_s · h_i^{(l)} + ā_i + b ``` where `W_n, W_s ∈ ℝ^{d_out × d_in}`, `W_e ∈ ℝ^{d_out × d_edge}`, and `b ∈ ℝ^{d_out}`. The full GraphSAGE model stacks 2 EdgeAwareSAGEConv layers with LayerNorm and Dropout: ``` h^(0) = x h^(1) = Dropout(ReLU(LayerNorm(SAGE_1(h^(0), edge_index, edge_attr)))) h^(2) = SAGE_2(h^(1), edge_index, edge_attr) ``` **Parameters:** With `d_in=26, d_hidden=128, d_out=64, d_edge=5`: Layer 1 has `26×128 + 26×128 + 5×128 + 128 = 7,424` params; Layer 2 has `128×64 + 128×64 + 5×64 + 64 = 16,768` params. Total: **~24.2K** (well under 500K). #### 3.6.2 GATv2SceneGraph Using GATv2Conv [32] with native edge feature support: ``` Layer 1: GATv2Conv(26 → 32, heads=4, concat=True, edge_dim=5) → output: 128 + LayerNorm(128) + ELU + Dropout(0.2) Layer 2: GATv2Conv(128 → 64, heads=1, concat=False, edge_dim=5) → output: 64 ``` GATv2 computes dynamic attention: ``` α_ij = softmax_j(a^T · LeakyReLU(W·[h_i ‖ h_j] + W_e·e_ij)) h_i' = Σ_j α_ij · W_v · h_j ``` **Parameters:** ~24.5K. #### 3.6.3 HybridGNN We combine both architectures with a learned gate: ``` h_sage = GraphSAGE(x, edge_index, edge_attr) h_gat = GATv2(x, edge_index, edge_attr) g = σ(W_gate · [h_sage ‖ h_gat]) h_out = g ⊙ h_sage + (1-g) ⊙ h_gat ``` where `W_gate ∈ ℝ^{64×128}` and `σ` is the element-wise sigmoid. This allows the model to automatically blend local aggregation (SAGE) and attention (GAT) per node. **Parameters:** ~57.3K (sum of SAGE + GAT + gate). All three models include a final output projection: `[Linear(64→64) → LayerNorm → ReLU]`, adding ~4.5K parameters. Total parameter counts including projection: **SAGE: 28.7K, GAT: 29.3K, Hybrid: 62.0K**. ### 3.7 Loss Functions #### 3.7.1 Edge-Weighted Cross-Entropy ``` L_CE^edge = (1/|V|) Σ_{(x,y) ∈ V} w(x,y) · CE(S(x,y), y*(x,y)) ``` where `V = {(x,y) : y*(x,y) ≠ ignore_index}` is the set of valid pixels, and: ``` w(x, y) = 1 + α_edge · C(x, y) ``` with `α_edge = 2.0`. Boundary pixels (high `C`) receive up to 3× the weight of interior pixels. #### 3.7.2 Boundary Loss (SBCB-Style) We compute boundary ground truth on-the-fly from segmentation labels using a Laplacian kernel: ``` B_gt = 𝟙(|L * y*| > 0.5) ``` where `L` is the discrete Laplacian `[[0,1,0],[1,-4,1],[0,1,0]]`. `B_gt` is then dilated with a `(2d+1)×(2d+1)` kernel (`d=2` pixels) to create a soft boundary region. The boundary loss is binary cross-entropy with positive class weighting: ``` L_boundary = (1/|Ω|) Σ_{(x,y)} w_+(x,y) · BCE(B_pred(x,y), B_gt(x,y)) ``` where `w_+ = min(|Ω|/|Ω_+|, 20)` compensates for the sparsity of boundary pixels (~5% of the image). #### 3.7.3 Combined Segmentation Loss ``` L_seg = L_CE^edge + λ · L_boundary ``` We set `λ = 0.4` based on SBCB's finding that boundary losses should be subordinate to the primary segmentation objective (their `β/α = 1/5 = 0.2`; we increase this slightly given our edge-weighted CE already provides boundary focus). #### 3.7.4 Relational Consistency Loss (GNN) For GNN training, we employ a contrastive loss that encourages similar refined features for same-class neighbors and dissimilar features for different-class neighbors: ``` L_rel = (1/|E|) Σ_{(i,j)∈E} [s_ij · ‖h_i' - h_j'‖² + (1-s_ij) · max(0, m - ‖h_i' - h_j'‖)²] ``` where `s_ij = 𝟙(class_i = class_j)` and `m = 1.0` is the margin. #### 3.7.5 Chamfer Distance (Evaluation) For evaluating 3D reconstruction quality: ``` CD(P, Q) = (1/|P|) Σ_{p∈P} min_{q∈Q} ‖p-q‖² + (1/|Q|) Σ_{q∈Q} min_{p∈P} ‖q-p‖² ``` ### 3.8 Stage 5: 3D Point Cloud Generation For each primitive `p_i`, we sample `n_s = 512` points on its canonical surface: - **Cuboid**: Uniform sampling on 6 faces (`n_s/6` per face) - **Cylinder**: 75% on the lateral surface, 12.5% on each cap - **Cone**: 70% on the conical surface (area-weighted via `√t` sampling), 30% on the base - **Plane**: Uniform sampling on the XZ plane with near-zero Y variance Sampled points are transformed via: `p_world = R(q_i) · (p_local ⊙ size_i) + centroid_i + ε`, where `R(q)` converts quaternion to rotation matrix and `ε ~ N(0, σ²I)` with `σ = 0.02` adds surface noise for realism. The output point cloud includes per-point metadata: `{position ∈ ℝ³, class_id, instance_id, primitive_type, color ∈ ℝ³, gnn_features ∈ ℝ^d}`. --- ## 4. Training Strategy We adopt a 4-phase training strategy that progressively introduces complexity: ### Phase 1: Segmentation Pretraining (30 epochs) - Train UNet with standard (non-edge-weighted) cross-entropy - Optimizer: AdamW, lr = 1×10⁻³, weight decay = 1×10⁻⁴ - Scheduler: Cosine annealing - Boundary head disabled - Purpose: Establish stable feature representations ### Phase 2: Edge-Weighted Fine-Tuning (15 epochs) - Switch to combined loss: `L_CE^edge + 0.4 · L_boundary` - Enable boundary head - Optimizer: AdamW, lr = 5×10⁻⁴ - Purpose: Sharpen boundary predictions ### Phase 3: GNN Training (50 epochs) - Train GNN on scene graphs extracted from Phase 2 segmentations - Relational consistency loss - Optimizer: Adam, lr = 1×10⁻³ - Segmentation weights frozen - Purpose: Learn spatial relationships ### Phase 4: End-to-End Fine-Tuning (10 epochs) - Joint training of all components - Freeze encoder BatchNorm layers (prevent distribution shift) - Optimizer: AdamW, lr = 1×10⁻⁴ - Combined loss: `L_seg + 0.1 · L_rel` - Gradient clipping: max_norm = 1.0 - Purpose: Harmonize all stages --- ## 5. Experimental Design ### 5.1 Datasets | Dataset | Images | Resolution | Classes | Use | |---|---|---|---|---| | **Cityscapes** [35] | 2,975/500/1,525 | 2048×1024 | 19 | Primary train/val/test | | **BDD100K** [36] | 7,000/1,000/2,000 | 1280×720 | 19 | Robustness evaluation | | **CARLA** [37] | Configurable | Configurable | Configurable | 3D GT supervision | Cityscapes provides fine pixel-level semantic annotations for urban street scenes. BDD100K offers diverse driving conditions (weather, time-of-day, scene type). CARLA provides synthetic data with exact 3D ground truth (depth maps, 3D bounding boxes, semantic point clouds) for quantitative 3D evaluation. ### 5.2 Evaluation Metrics | Metric | Formula | Target | |---|---|---| | **3D IoU** | `V_intersect / V_union` (AABB) | ~0.68 | | **Centroid L2** | `(1/N) Σ min_j ‖c_i^pred - c_j^gt‖₂` | ~0.49 m | | **Edge Accuracy** | F1 of predicted vs GT edges | ~78% | | **Chamfer Distance** | Bidirectional nearest-neighbor | ~0.041 | | **Boundary IoU** | IoU restricted to boundary pixels | +15% over baseline | | **mIoU** | Mean Intersection-over-Union | Competitive | | **FPS** | Frames per second (RTX 3090) | ≥15 | | **Parameters** | GNN module parameter count | <500K | ### 5.3 Baselines We compare against: 1. **Non-edge baseline**: Same UNet without edge weighting or boundary loss 2. **MonoScene** [5]: State-of-the-art monocular SSC 3. **Depth Anything V2** [28] + back-projection: Direct depth-to-pointcloud lifting 4. **STDC-Seg** [11] + naive 3D: Real-time segmentation without relational reasoning ### 5.4 Implementation Details - **Framework**: PyTorch 2.0+ with PyTorch Geometric 2.4+ - **Hardware**: Training on NVIDIA A100 (80GB); inference benchmarked on RTX 3090 - **Input resolution**: 512×1024 (Cityscapes crops) and 256×512 (fast mode) - **Augmentation**: Random horizontal flip, scale jitter [0.5, 2.0], color jitter, Gaussian blur - **Batch size**: 8 (Phase 1–2), variable (Phase 3), 4 (Phase 4) - **Weight initialization**: Kaiming normal for convolutions, Xavier uniform for GNN layers --- ## 6. Ablation Studies We design four ablation experiments to validate key design choices: ### 6.1 Edge Loss Weight (λ) We sweep `λ ∈ {0.0, 0.1, 0.2, 0.4, 0.6, 0.8, 1.0}` and measure: - Boundary IoU improvement over `λ=0` baseline - Overall mIoU (to detect over-regularization) - Training stability (loss curves) **Hypothesis**: `λ ∈ [0.3, 0.5]` maximizes boundary IoU while maintaining overall mIoU. ### 6.2 GNN Architecture We compare EdgeAwareSAGE, GATv2, and Hybrid on: - Relational consistency (same-class feature similarity) - Edge prediction accuracy - Inference latency - Parameter count **Hypothesis**: Hybrid outperforms individual models; GATv2 provides better attention visualization for interpretability; SAGE is fastest. ### 6.3 Edge Detection Method We compare Sobel vs. differentiable Canny: - Edge precision/recall against GT boundaries - Downstream segmentation quality - Computational overhead **Hypothesis**: Canny provides sharper edges (double threshold), while Sobel is faster and more differentiable-friendly. ### 6.4 Points per Primitive We sweep `n_s ∈ {128, 256, 512, 1024, 2048}` and measure: - Chamfer distance to ground truth - Visual quality - Generation time **Hypothesis**: Diminishing returns beyond 512 points; 256 is sufficient for most downstream tasks. --- ## 7. Preliminary Results We report preliminary results from synthetic traffic scene testing (no dataset training—random weights with synthetic segmentation inputs): ### 7.1 Parameter Budget Verification | GNN Variant | GNN Params | Under 500K | Total Pipeline | |---|---|---|---| | EdgeAwareSAGE | 28,736 | ✓ | 4.36M | | GATv2 | 29,312 | ✓ | 4.36M | | Hybrid | 62,016 | ✓ | 4.39M | All variants satisfy the <500K GNN constraint by a factor of >8×, providing substantial headroom for architecture scaling if needed. ### 7.2 Per-Stage Latency (CPU, 128×256) | Stage | Module | Latency | |---|---|---| | Stage 1 | Input Augmentation | 1.3 ms | | Stage 2 | UNet Segmentation | 77.2 ms | | Stage 3 | Primitive Extraction | 30.5 ms | | Stage 4 | GNN Refinement | <0.1 ms | | Stage 5 | Point Cloud Generation | 0.6 ms | | **Total** | | **~110 ms** | The segmentation UNet dominates latency (70%). On GPU, we expect ~10× speedup for the UNet, bringing total pipeline latency to ~15–20ms (50–67 FPS), comfortably exceeding the 15 FPS target. ### 7.3 Pipeline Validation The end-to-end pipeline successfully: - Produces 5-channel augmented tensors with valid value ranges - Generates K-class segmentation maps with boundary predictions - Extracts 4–8 primitives per synthetic scene with correct types - Constructs scene graphs with 20+ edges per scene - Refines node features via all three GNN variants - Generates 2K–20K colored, labeled point clouds - Exports valid PLY files for visualization --- ## 8. Discussion ### 8.1 Advantages of the Proposed Approach **Compactness.** Where MonoScene operates on 256×256×32 = 2M voxels, Traffic3D represents a scene with 4–64 primitives (~100× compression). This enables faster downstream reasoning and reduced memory footprint. **Interpretability.** Each primitive has a semantic type, centroid, size, and orientation—directly corresponding to human-understandable scene elements. This contrasts with opaque voxel occupancy grids or latent feature volumes. **Modularity.** Each stage can be independently upgraded (e.g., replacing UNet with STDC-Seg for higher speed, or substituting the depth prior with Depth Anything V2 for metric depth) without modifying other stages. **Zero-overhead boundary supervision.** The SBCB-style boundary head is discarded at inference, providing boundary quality improvement at no computational cost—a property not shared by SegFix or InverseForm, which add post-processing modules. ### 8.2 Limitations and Mitigations **Depth approximation.** Our vertical positional encoding is a coarse depth proxy that fails for objects at unusual heights (elevated roads, bridges) or non-standard camera angles. *Mitigation*: Integrate Depth Anything V2-Small (24.8M params) as an optional depth backbone. **Primitive expressiveness.** Simple cuboids/cylinders cannot capture complex geometries (curved buildings, non-standard vehicles). *Mitigation*: Replace with differentiable superquadrics [15, 25] in future work, adding the squareness parameters `(ε₁, ε₂)`. **Disconnected training.** The 4-phase strategy trains stages sequentially, potentially missing joint optimization opportunities. *Mitigation*: Phase 4 (end-to-end) partially addresses this, but a fully differentiable primitive extraction step would enable true end-to-end gradient flow. **Connected component bottleneck.** The scipy-based connected component analysis runs on CPU and becomes the latency bottleneck for complex scenes. *Mitigation*: Replace with GPU-accelerated connected components (e.g., cuCOMP from RAPIDS) or learned instance segmentation. ### 8.3 Broader Impact Traffic3D could enable low-cost 3D scene understanding for autonomous vehicles in developing markets where LiDAR is prohibitively expensive, facilitate urban planning through automated 3D mapping from dashcam footage, and support augmented reality navigation overlays. We note that any vision-based driving system should be validated for fairness across demographics and lighting conditions before deployment. --- ## 9. Future Work and Extensions ### 9.1 Short-Term Extensions 1. **Metric depth integration**: Replace positional encoding with Depth Anything V2-Small for accurate metric depth, enabling true-scale 3D reconstruction. 2. **Temporal consistency**: Extend the scene graph with temporal edges between corresponding primitives across frames, using the GNN to enforce temporal smoothness. 3. **Instance segmentation**: Replace connected components with a lightweight instance segmentation head (e.g., SOLO or mask embedding) for better object separation. ### 9.2 Medium-Term Research Directions 4. **Differentiable primitive fitting**: Make the primitive extraction step fully differentiable, enabling end-to-end training from image to 3D reconstruction with 3D supervision from CARLA. 5. **Superquadric upgrade**: Replace simple primitives with learned superquadrics following SuperOcc [15], adding squareness and opacity parameters for richer geometry. 6. **Multi-scale GNN**: Implement hierarchical message passing with local (object-level) and global (scene-level) graphs for better long-range reasoning. ### 9.3 Long-Term Vision 7. **Self-supervised pre-training**: Learn scene graph representations from unlabeled driving video using contrastive learning between augmented views of the same scene. 8. **Language-grounded scene graphs**: Integrate CLIP-style [38] embeddings into node features for natural language querying of 3D scenes ("find the red car near the crosswalk"). 9. **Dynamic scene modeling**: Estimate per-primitive velocities from temporal scene graphs, enabling prediction of future scene states for motion planning. --- ## 10. Conclusion We have presented Traffic3D, a modular pipeline for monocular 3D traffic scene reconstruction that bridges the gap between 2D semantic segmentation and 3D scene understanding. Through the combination of edge-aware augmentation, boundary-conditioned supervision, PCA-based primitive fitting, lightweight GNN relational refinement, and surface-sampled point cloud generation, our approach provides an interpretable, compact, and efficient alternative to dense voxel-based methods. Key technical contributions include: (i) a novel edge-weighted loss function with zero-overhead boundary supervision, (ii) the EdgeAwareSAGEConv layer that extends GraphSAGE with edge feature injection, (iii) a Hybrid GNN with learned gating between aggregation and attention paradigms, and (iv) a complete, modular codebase with 4-phase training, comprehensive evaluation metrics, and an ablation study framework. Preliminary validation confirms that all three GNN variants satisfy the <500K parameter constraint (28.7K–62.0K), the full pipeline totals 4.36M parameters, and the system successfully produces semantically labeled 3D point clouds from single RGB images. We anticipate that with training on Cityscapes/CARLA data, Traffic3D will achieve competitive metrics (3D IoU ~0.68, Centroid L2 ~0.49m, Chamfer distance ~0.041) while maintaining ≥15 FPS inference on an RTX 3090. The modular architecture ensures that each stage can be independently upgraded as the field advances—making Traffic3D not just a static system but a living framework for monocular 3D scene understanding research. --- ## References [1] C. R. Qi, H. Su, K. Mo, and L. J. Guibas. "PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation." CVPR, 2017. [2] C. R. Qi, L. Yi, H. Su, and L. J. Guibas. "PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space." NeurIPS, 2017. [3] J. Philion and S. Fidler. "Lift, Splat, Shoot: Encoding Images from Arbitrary Camera Rigs by Implicitly Unprojecting to 3D." ECCV, 2020. [4] Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Q. Yu, and J. Dai. "BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers." ECCV, 2022. [5] A.-Q. Cao and R. de Charette. "MonoScene: Monocular 3D Semantic Scene Completion." CVPR, 2022. arXiv:2112.00726. [6] Y. Li, Z. Yu, C. Choy, C. Xiao, J. M. Alvarez, S. Fidler, C. Feng, and A. Anandkumar. "VoxFormer: Sparse Voxel Transformer for Camera-based 3D Semantic Scene Completion." CVPR, 2023. arXiv:2302.12251. [7] R. Zellers, M. Yatskar, S. Thomson, and Y. Choi. "Neural Motifs: Scene Graph Parsing with Global Context." CVPR, 2018. [8] K. Tang, Y. Niu, J. Huang, J. Shi, and H. Zhang. "Unbiased Scene Graph Generation from Biased Training." CVPR, 2020. [9] X. Lin, C. Ding, J. Zeng, and D. Tao. "GPS-Net: Graph Property Sensing Network for Scene Graph Generation." CVPR, 2020. [10] H. Araki, J. Yao, and G. Shakhnarovich. "Boosting Semantic Segmentation with Semantic Boundaries." 2023. arXiv:2304.09427. [11] M. Fan, S. Lai, J. Huang, X. Wei, Z. Chai, J. Luo, and X. Wei. "Rethinking BiSeNet For Real-time Semantic Segmentation." CVPR, 2021. arXiv:2104.13188. [12] S. Song, F. Yu, A. Zeng, A. X. Chang, M. Savva, and T. Funkhouser. "Semantic Scene Completion from a Single Depth Image." CVPR, 2017. [13] Y. Tong, F. Li, H. Wang, et al. "Scene as Occupancy." ICCV, 2023. [14] Y. Wei, L. Zhao, W. Zheng, Z. Zhu, J. Zhou, and J. Lu. "SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous Driving." ICCV, 2023. [15] Y. Chen et al. "SuperOcc: Toward Cohesive Temporal Modeling for Superquadric-based 3D Occupancy Prediction." 2025. arXiv:2601.15644. [16] C. Yu, J. Wang, C. Peng, C. Gao, G. Yu, and N. Sang. "BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation." ECCV, 2018. [17] O. Ronneberger, P. Fischer, and T. Brox. "U-Net: Convolutional Networks for Biomedical Image Segmentation." MICCAI, 2015. [18] Y. Yuan, X. Chen, and J. Wang. "SegFix: Model-Agnostic Boundary Refinement for Segmentation." ECCV, 2020. [19] S. Borse, Y. Wang, Y. Zhang, and F. Porikli. "InverseForm: A Loss Function for Structured Boundary-Aware Segmentation." CVPR, 2021. [20] A. Liu et al. "BPKD: Boundary Privileged Knowledge Distillation For Semantic Segmentation." 2023. arXiv:2306.08075. [21] M. Music, J. Delplanque, M. Ianetta, and P. Music. "REACT: Real-time Efficiency and Accuracy Compromise for Tradeoffs in Scene Graph Generation." 2024. arXiv:2405.16116. [22] T. Li et al. "Graph-based Topology Reasoning for Driving Scenes." 2023. arXiv:2304.05277. [23] M. Yang, D. Hu, M. Ding, and C. Hao. "Joint Generative Modeling of Scene Graphs and Images via Diffusion Models." 2024. arXiv:2401.01130. [24] Y. Liu et al. "Controllable 3D Outdoor Scene Generation via Scene Graphs." 2025. arXiv:2503.07152. [25] T. Monnier, J. Austin, A. Kanazawa, A. Efros, and M. Aubry. "Differentiable Blocks World: Qualitative 3D Decomposition by Rendering Primitives." NeurIPS, 2023. arXiv:2307.05473. [26] Y. Liu and G. S. Chirikjian. "Marching-Primitives: Shape Abstraction from Signed Distance Function." 2023. arXiv:2303.13190. [27] "SuperDec: 3D Scene Decomposition with Superquadric Primitives." 2025. arXiv:2504.00992. [28] L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao. "Depth Anything V2." 2024. arXiv:2406.09414. [29] S. Bochkovskiy et al. "Depth Pro: Sharp Monocular Metric Depth in Less Than a Second." Apple ML, 2024. arXiv:2410.02073. [30] "AnyDepth: Depth Estimation Made Easy." 2025. arXiv:2601.02760. [31] W. L. Hamilton, R. Ying, and J. Leskovec. "Inductive Representation Learning on Large Graphs." NeurIPS, 2017. arXiv:1706.02216. [32] S. Brody, U. Alon, and E. Yahav. "How Attentive are Graph Attention Networks?" ICLR, 2022. arXiv:2105.14491. [33] P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio. "Graph Attention Networks." ICLR, 2018. [34] C. Gao, H. Bai, S. Zhou, and J. Ran. "EGNAS: Edge Feature Graph Neural Architecture Search." 2021. arXiv:2109.01356. [35] M. Cordts, M. Omran, S. Ramos, et al. "The Cityscapes Dataset for Semantic Urban Scene Understanding." CVPR, 2016. [36] F. Yu, H. Chen, X. Wang, et al. "BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning." CVPR, 2020. [37] A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun. "CARLA: An Open Urban Driving Simulator." CoRL, 2017. [38] A. Radford, J. W. Kim, C. Hallacy, et al. "Learning Transferable Visual Models From Natural Language Supervision." ICML, 2021. --- ## Appendix A: Complete System Architecture ``` ┌─────────────────────────────────────────────────────────────────────────┐ │ TRAFFIC3D PIPELINE │ │ │ │ Input: RGB Image I ∈ ℝ^{H×W×3} │ │ │ │ ┌─────────────────────────────────────────────────────────────────┐ │ │ │ STAGE 1: INPUT AUGMENTATION │ │ │ │ │ │ │ │ I ──→ [Normalize] ──→ RGB ∈ [0,1] │ │ │ │ [Luminance] ──→ I_gray │ │ │ │ [Sobel/Canny] ──→ Edge Map C ∈ [0,1] │ │ │ │ [Positional] ──→ P(x,y) = y/H │ │ │ │ [Concat] ──→ F ∈ ℝ^{H×W×5} │ │ │ └────────────────────────────┬────────────────────────────────────┘ │ │ │ │ │ ┌────────────────────────────▼────────────────────────────────────┐ │ │ │ STAGE 2: EDGE-WEIGHTED SEGMENTATION │ │ │ │ │ │ │ │ F ──→ [UNet Encoder: 5→32→64→128→256→256] │ │ │ │ ──→ [UNet Decoder: 256→128→64→32→32] │ │ │ │ ──→ [1×1 Conv] ──→ S ∈ ℝ^{H×W×K} │ │ │ │ ──→ [Edge Weight] ──→ S' = S·(1+α·C) │ │ │ │ ──→ [Boundary Head]* ──→ B_pred ∈ [0,1] │ │ │ │ (*training only) │ │ │ │ Loss: L_CE^edge + λ·L_boundary │ │ │ └────────────────────────────┬────────────────────────────────────┘ │ │ │ seg = argmax(S') │ │ ┌────────────────────────────▼────────────────────────────────────┐ │ │ │ STAGE 3: PRIMITIVE EXTRACTION + SCENE GRAPH │ │ │ │ │ │ │ │ seg ──→ [Connected Components per class] │ │ │ │ ──→ [Depth Prior: Z = d·(1-y/H) + Z_min] │ │ │ │ ──→ [Pixel→3D Projection] │ │ │ │ ──→ [PCA Fitting] ──→ Primitives {type,c,s,q} │ │ │ │ ──→ [Graph Build] ──→ G = (V, E, X, A) │ │ │ │ Nodes: [emb(16) ‖ c(3) ‖ s(3) ‖ q(4)] = 26D │ │ │ │ Edges: dist < 5m, features: [d,adj,Δpos] = 5D │ │ │ └────────────────────────────┬────────────────────────────────────┘ │ │ │ │ │ ┌────────────────────────────▼────────────────────────────────────┐ │ │ │ STAGE 4: GNN RELATIONAL REFINEMENT │ │ │ │ │ │ │ │ Option A: EdgeAwareSAGE (28.7K params) │ │ │ │ m_ij = W_n·h_j + W_e·e_ij; h_i' = W_s·h_i + MEAN(m) │ │ │ │ │ │ │ │ Option B: GATv2 (29.3K params) │ │ │ │ α_ij = softmax(a^T·LeakyReLU(W·[h_i‖h_j]+W_e·e_ij)) │ │ │ │ h_i' = Σ α_ij · W_v · h_j │ │ │ │ │ │ │ │ Option C: Hybrid (62.0K params) │ │ │ │ g = σ(W·[h_sage ‖ h_gat]); h' = g⊙h_sage + (1-g)⊙h_gat │ │ │ │ │ │ │ │ All: 2 layers, hidden=128, out=64, LayerNorm, Dropout(0.2) │ │ │ │ + Output projection: Linear→LayerNorm→ReLU │ │ │ └────────────────────────────┬────────────────────────────────────┘ │ │ │ │ │ ┌────────────────────────────▼────────────────────────────────────┐ │ │ │ STAGE 5: 3D POINT CLOUD GENERATION │ │ │ │ │ │ │ │ For each primitive p_i: │ │ │ │ [Sample ~512 pts on canonical surface] │ │ │ │ [Scale by size_i] │ │ │ │ [Rotate by R(q_i)] │ │ │ │ [Translate by centroid_i] │ │ │ │ [Add noise ε ~ N(0, 0.02²)] │ │ │ │ [Assign: class, instance, color, GNN features] │ │ │ │ │ │ │ │ Output: P = {(xyz, class, instance, prim_type, color, feat)} │ │ │ │ 2K–20K points, exportable as PLY │ │ │ └─────────────────────────────────────────────────────────────────┘ │ │ │ │ Output: Segmentation Map + Scene Graph + 3D Point Cloud │ └─────────────────────────────────────────────────────────────────────────┘ ``` ## Appendix B: Cityscapes Class-to-Primitive Mapping | ID | Class | Primitive | Depth Scale | Rationale | |---|---|---|---|---| | 0 | road | Plane | 0.1 | Flat ground surface | | 1 | sidewalk | Plane | 0.1 | Flat ground surface | | 2 | building | Cuboid | 2.0 | Rectangular structures | | 3 | wall | Cuboid | 0.5 | Rectangular structures | | 4 | fence | Cuboid | 0.3 | Rectangular structures | | 5 | pole | Cylinder | 0.15 | Thin vertical objects | | 6 | traffic light | Cuboid | 0.5 | Boxy objects | | 7 | traffic sign | Cuboid | 0.3 | Flat rectangular objects | | 8 | vegetation | Cone | 1.5 | Triangular canopy shape | | 9 | terrain | Plane | 0.1 | Flat ground surface | | 10 | sky | Plane | 0.1 | Infinite flat backdrop | | 11 | person | Cylinder | 0.5 | Upright cylindrical form | | 12 | rider | Cylinder | 0.5 | Upright cylindrical form | | 13 | car | Cuboid | 2.0 | Rectangular vehicle body | | 14 | truck | Cuboid | 3.0 | Large rectangular body | | 15 | bus | Cuboid | 4.0 | Large rectangular body | | 16 | train | Cuboid | 1.0 | Long rectangular body | | 17 | motorcycle | Cuboid | 1.2 | Small rectangular body | | 18 | bicycle | Cuboid | 0.8 | Small rectangular body | ## Appendix C: Reproducibility The complete implementation is publicly available at: **Repository**: [https://huggingface.co/HugMaster2002/traffic3d-monocular-reconstruction](https://huggingface.co/HugMaster2002/traffic3d-monocular-reconstruction) The repository includes: - All 5 pipeline stages as modular PyTorch modules - 3 GNN variants (EdgeAwareSAGE, GATv2, Hybrid) - 5 loss functions (EdgeWeightedCE, BoundaryLoss, CombinedSeg, RelationalConsistency, Chamfer) - 4-phase training pipeline - Complete evaluation suite with 6 metrics - Ablation study framework with 4 ablation dimensions - End-to-end test suite validating all components - PLY point cloud export for visualization **Dependencies**: PyTorch ≥2.0, PyTorch Geometric ≥2.4, scipy, scikit-learn, numpy. --- *End of Proposal*