The US military wants autonomous tank fire control, but the bottleneck isn't the AI model—it's the data
What determines whether tanks and autonomous vehicles can fight isn't model architecture—it's whether anyone is willing to keep doing the boring, unglamorous work of data collection, labeling, and governance that nobody wants to do.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
"AI-ready data" is a false premise
Bharat opens by demolishing the notion of ‘AI-ready data’ as an objective: data will never truly be ready for AI—it inherently comes with various ‘quirks’. What should be determined first is AI's concrete use case, because the use case determines what kind of data is needed and what quality threshold matters. Many teams get the effort allocation backward from the start, chasing what seems like a clean, perfect dataset instead of first clarifying what problem they're solving.
— Bharat PatelAutonomous tank vision requires purpose-built data collection systems
When building computer vision models for tanks at an Army research lab, the team discovered there was no existing tank dataset. So they built their own data collection box, mounted beside the sensors, to specifically capture second-gen FLIR imagery. But it covered only one sensor type—data from the tank's other sensors wasn't being collected. True autonomous capability requires multi-modal data fusion, not single-sensor inference, a limitation they didn't recognize until later.
— Bharat PatelThree years of accumulated battle footage enabled Ukrainian autonomous weapons
The autonomous systems Ukraine operates today didn't start that way. Early operators controlled devices with joysticks; video footage got systematically recovered and accumulated—a process many participants didn't even realize was happening. It's that years-long accumulation of video data that later enabled training for target recognition and autonomous capability. Meanwhile, the US Department of Defense lacks a systematic, continuous data collection strategy; exercise-generated data typically gets deleted afterward or warehoused without context, never reused.
— Bharat PatelGovernment controls data, industry executes—roughly ninety percent of the work
Bharat's ideal division of labor: data governance authority stays with government; government also designs the data pipeline, processes, and governance standards. But execution should be outsourced as much as possible—data annotation, testing and evaluation, deployment, model training. Industry should handle roughly 90 percent of the work. The key: government must structure these tasks as modular contracts that allow continuous iteration with industry without lock-in to any single contractor.
— Bharat PatelPentagon speed comes from putting top talent on priorities
On the Pentagon's PAE/CPE restructuring push over the past two years, Bharat argues that organizational reshuffling misses the core issue. What matters is holding project managers accountable for capability delivery and giving them matched resources. He's blunt: he's tired of hearing ‘we need to accelerate’ while nothing actually accelerates. The real mechanism for speed is simple: concentrate your fastest, most motivated people on high-priority projects. They can ‘move mountains’.
— Bharat PatelModels fail on simple adversarial examples they never saw in training
Data poisoning takes two forms. One: in open-source data ecosystems, actors deliberately inject mislabeled or biased data—the model absorbs these errors during training. Two: real-world adversarial examples like spray-painting a target—the model never saw painted targets in training, so it fails. Even at the electromagnetic spectrum level, adversaries can use a sensor-laden vehicle to fabricate signal signatures that mimic ‘large-scale military operations’.
— Bharat PatelCombat validates what testing cannot: enemies and circumstances both have a vote
Testing can't stretch indefinitely—models ultimately get validated in combat, where ‘the enemy has a vote, the environment has a vote, and the soldiers using the model have a vote’—no amount of testing gives absolute coverage. The realistic approach: work with intelligence to turn known adversarial tactics (like painting targets) into synthetic data, inject it into training and test sets, so models learn to recognize and handle these interruptions. Stop chasing perfect test coverage.
— Bharat PatelDecoupling AI from software certification lets models update faster than platforms
Current military systems certify software as a monolithic whole under RMF, meaning updating a single model requires recertifying the entire software stack. Before leaving, Bharat pushed to decoupling AI—data and models—from the software stack entirely. That way, model iterations wouldn't have to drag the whole software package through re-certification, and autonomous capability could be pushed to field equipment as fast and frequently as OTA updates.
— Bharat PatelIn their own words · checked verbatim
In my opinion, data will never be AI ready. Data will always have its weirdness about it.
Bharat Patel2:07
So the team actually had to build a data collection box. They had to build a data collection box. They had to put it right next to the sensor that they wanted the data from.
Bharat Patel4:18
So there is a very long period of time where there is an active data collection strategy that was happening in beknownst to the people that were there.
Bharat Patel9:42
I'm so tired of like, you know, we got to go fast. We're going to go fast. And, you know, nobody ever really goes fast.
Bharat Patel21:51
it's still fairly easy to put camouflage over something or just spray paint graffiti over a target and it no longer becomes a target because that model has never seen a target with graffiti on it.
Bharat Patel30:25
the enemy is going to have a vote, the environment is going to have a vote, the warfighter that's using the model is going to have a vote.
Bharat Patel31:27
once the, the shiny toy of products is over, the department is going to realize you need to actually a qualified integration team to make everything work.
Bharat Patel38:50
Figures
| Project Maven launch | 2016 | 2:07 |
| Bharat entered Army procurement | 2007 | 16:11 |
| Titan synthetic data exploration started | ~2022 | 25:07 |
| Ideal industry proportion of data pipeline work | ~90% | 15:00 |
| Model acceptable performance threshold | 80% | 31:27 |
| Bharat's time at Accenture | ~6 months | 38:50 |
Glossary
- Project Maven
- US Department of Defense computer vision initiative launched in 2016 to identify targets in drone footage
- CDAO
- Chief Digital and AI Officer of the Department of Defense; oversees AI data standards, governance, and infrastructure
- RMF
- Risk Management Framework—military cybersecurity certification process that governs software system approval
- PAE/CPE
- Program Executive/Capability Program Executive—new Pentagon procurement management roles coordinating cross-project requirements
- second-gen FLIR
- Second-generation Forward-Looking Infrared imaging—thermal sensor used in tanks and armor for night vision and target identification
- SBIR Direct to Phase II
- US government program allowing small businesses to skip research phase 1 and receive phase 2 research funding directly
How to listen
Founders and investors interested in defense tech, AI militarization, and government procurement; engineers who want to understand the real constraints of autonomous weapons, not the hype.
After 39:54, there's a section on Accenture recruiting and self-promotion with low information density; you can skip it.