AV data annotation
A self-driving car has to recognize every vehicle and person around it. We label what it sees.
Self-driving cars only work because someone labels what they see — pedestrians, brake lights, curbs, and the edge cases models fail on. DesiCrew annotates that data in 2D and 3D, across eight camera views and every frame, so the model learns from clean, consistent ground truth.
5 min read
Self-driving cars only work because someone labels what they see — pedestrians, brake lights, curbs, and the edge cases that models fail on. DesiCrew annotates that data for the broader AV ecosystem in 2D and 3D, across eight camera views and every frame — so the model learns from clean, consistent ground truth.
The model can't tell a parked car from a moving one, or a pedestrian from a rider, unless a person makes that call first. The client's sensors and automation capture the scene. DesiCrew's people decide what each object is, where its edges sit, and whether it's the same object one frame later. So the AI learns from ground truth it can trust — and the same object is never counted twice.
01 — The Mandate
The client runs three annotation processes; DesiCrew runs the vehicle-and-human labeling tier that trains the autonomous-driving model, delivered on the client's Deepen platform.
Scale. Each file runs to around 48 frames, and every frame is covered by eight camera angles — front, sides, the four corners, and a bird's-eye view — checked in both 2D and 3D. Coverage. This tier covers vehicles and people only; the client completed the earlier environment-mapping tiers before this stage, so poles, buildings and other fixed objects sit there, not here. Three lines of work. Annotating each object, tracking it across every frame, and quality-checking the output before it is accepted. Built for continuity. Headcount is split 50/50 across two sites so the work carries a full business-continuity backup.
02 — The Challenge
Drawing the box is the easy part. Agreeing where its edge sits is the hard part. Get the boundary of a vehicle slightly wrong and the model learns the wrong shape. Most errors on this work aren't misreadings — they are disagreements about exactly where an object begins and ends. That is a matter of perception, and it is why the work resists full automation.
A single object is not a single click. One vehicle carries a category, an occlusion grade, a state — parked, in-parking or driving — a protruding-object flag for things like mirrors, and a box that has to stay aligned as the object moves. All of it has to hold across ~48 frames and all eight cameras. The 2D and 3D have to agree. The same object must be labeled in both the flat camera image and the LiDAR point cloud, with a matching ID synced across the two — if they drift apart, the model is fed duplicates instead of one clean object. The rules are situational. The rulebook is the same across vehicle types — only the box dimensions change; occlusion is graded by what the view actually shows, from none to partial to full. The inputs are layered. A flat camera feed plus a moving 3D point cloud, eight angles per scene, and no scene context — the city the data came from is never exposed to the team.
03 — The Approach
DesiCrew labels by hand where judgment is required, checks a deliberate sample, and stays continuously calibrated to the client's eye.
Judgment on every object. Each object is tracked by a person across all frames and all eight cameras, with 2D and 3D kept in sync — so properties only change when the scene actually changes. Sampled quality control. After annotation, work goes to human QC: roughly 15–20% of frames, and around 30–40% of objects within a file, with IDs picked at random rather than in sequence. Quality built in up front. The team aims to land 90%+ quality at the annotation stage, so QC spends fewer hours per file and fewer errors travel downstream. Calibrated to the client. Because the hard errors are subjective, DesiCrew and the client run calibration sessions to get both sides seeing an object's edge the same way. Trained on the client's platform. New joiners are onboarded inside the platform's own tool and workspace — LiDAR reading taught from scratch — over about 15 days.
04 — The Outcome
One clean object, not two. Every vehicle and person is matched across 2D and 3D with a single synced ID, so the model learns from one consistent version of each object — and the training data stays clean as the volume scales.
Quality caught before it ships. A deliberate QC sample and a 90%+ first-pass target keep errors from reaching the model, with the residual geometry errors narrowing through calibration. Coverage that doesn't stop. A 50/50 split across two sites means the work has a built-in backup rather than a single point of failure — so delivery holds even when one site can't.
05 — Why the relationship holds
This is a young program — live for about five months — but it sits on a longer runway: the client signed the master agreement roughly two years before the work began. Enterprise clients take time to come on board, and once they do, they tend to stay.
Delivery runs across two sites on a 50/50 split, so continuity is built in rather than bolted on. And the relationship is kept aligned by design — regular calibration with the client's QC team keeps both sides working to the same standard as the work evolves.
About DesiCrew
DesiCrew is an applied-intelligence company — the human and technology layer that makes AI systems and enterprise operations work reliably and grow at scale, refined in production since 2007. IIT Madras incubated · Everest Group PEAK Matrix 2024 · Great Place to Work.
Let's put intelligence to work.
If your AI keeps breaking when it leaves the lab, or your operations are carrying weight AI should be taking off — let's build together.