The editorial argues the headline-grabbing 'AI beats human pilot' framing misses the point. The actual significance is that a reinforcement learning policy trained purely in simulation generalized well enough to a real F-16 that test pilots signed off on flying it in adversarial 9G engagements — sim-to-real is the unsolved problem across robotics, autonomous driving, and industrial control, and this is an aggressive demonstration of it working in a life-or-death domain.
DARPA frames ACE explicitly as human-machine teaming: the human pilot retains strategic control while the agent handles within-visual-range dogfighting, which operates at timescales too fast for deliberate human decision-making. A safety pilot with full override authority sat in the front seat throughout, and DARPA emphasizes that override was never triggered — presenting the program as incremental trust-building rather than a step toward lethal autonomy.
By submitting the DARPA release verbatim to a technical audience, the submitter surfaces the official framing that this is about human-machine teaming and trust calibration on the X-62A VISTA testbed rather than an operational combat capability.
DARPA and the U.S. Air Force Test Pilot School announced that the Air Combat Evolution (ACE) program has flown an AI agent in live, within-visual-range engagements against a human-piloted F-16. The airframe is the X-62A VISTA — a modified F-16D at Edwards Air Force Base whose flight control computer can be swapped out to let external software fly the jet. A safety pilot sat in the front seat with full override authority. According to the release, that override was never triggered during the combat maneuvers.
The agent was trained entirely in simulation using reinforcement learning, then deployed onto the real aircraft — the sim-to-real gap was closed well enough that a machine-learned policy pulled 9G turns against a human adversary without a human touching the stick. The tests ran multiple sorties across 2024 and 2025, progressing from defensive setups to offensive nose-on engagements. The program builds on the 2020 AlphaDogfight Trials, where a Heron Systems agent beat a Weapons School graduate 5-0 in simulation. What's new here is that the same class of agent is now flying steel, not pixels.
DARPA is careful to frame this as a trust-building exercise rather than a weapons program. ACE's stated goal is human-machine teaming — the pilot manages strategy, the agent handles the within-visual-range fight, which is the part of air combat that happens too fast for deliberate human decision-making anyway.
The headline everyone will run is "AI beats human pilot." That's the wrong headline. The interesting result is that a policy trained in a simulator generalized to a real airframe well enough that qualified test pilots signed off on flying it in an adversarial scenario. Sim-to-real is the unsolved problem in almost every applied RL domain — robotics, autonomous driving, industrial control — and this is one of the more aggressive demonstrations of it working under conditions where being wrong kills people.
Compare this to the state of autonomous driving. Waymo has spent roughly 15 years and untold billions to get robotaxis working in geofenced sunbelt cities at 45 mph. DARPA is claiming an ML agent can dogfight an F-16 at Mach speeds. The gap is real but not as wide as it looks: air combat has fewer edge cases than a suburban intersection (no pedestrians, no cyclists, no construction cones), the environment is better instrumented, and the failure mode is "eject" rather than "sued into oblivion." But the underlying stack — perception, state estimation, policy inference, control — is recognizably the same problem.
What's genuinely novel is the safety architecture: the agent runs on the VISTA's replaceable flight-control computer, with a hardware kill switch and a human in the loop who can revert to standard F-16 flight laws in milliseconds. That's a design pattern worth studying if you're building any autonomy system in a regulated domain. The agent doesn't need to be provably safe in the formal-methods sense; it needs to be observable, interruptible, and bounded by a simpler system that is provably safe. This is closer to how Tesla ships FSD than how the FAA certifies avionics, and the fact that the Air Force is willing to fly it says something about where the aerospace safety culture is headed.
The community response has been mixed in the way you'd expect. Some ex-military pilots on the thread argue this is a demo, not a doctrine — that BVR (beyond visual range) engagements dominate modern air combat and WVR dogfighting is a solved problem you avoid, not win. Others point out that unmanned platforms don't need to win dogfights against F-16s; they need to be cheap enough that losing ten of them for one manned adversary is a good trade. The CCA (Collaborative Combat Aircraft) program the Air Force is funding is built exactly on that thesis.
Most readers of this aren't building fighter jets. But there are three transferable lessons.
First, sim-to-real is getting cheaper and more reliable. If you're working on any robotics or control problem, the ACE result is another data point that training in a high-fidelity simulator and deploying to hardware is a viable path — you don't need the millions of hours of real-world data that the AV industry assumed were table stakes. Tooling like Isaac Sim, MuJoCo, and Genesis has closed enough of the reality gap that domain randomization plus modest fine-tuning works for a lot of tasks. If your team has been assuming "we need real-world data first," that assumption is worth re-testing every six months.
Second, the human-in-the-loop pattern is the design that ships. Fully autonomous is a research goal; supervised autonomy is the product. The reason the F-16 result exists is not that the AI is trustworthy — it's that the safety architecture makes trustworthiness irrelevant to whether you can fly the mission. That's the same reason Copilot ships and fully autonomous coding agents don't, and the same reason radiology AI is deployed as a second reader rather than a replacement. If you're designing an autonomy product, spend your engineering budget on the override, the observability, and the bounded operating envelope — not on making the agent perfect.
Third, defense contracting is quietly becoming a serious ML employer. Anduril, Shield AI, Palantir, and the primes are hiring RL and robotics engineers at rates that compete with FAANG, and the work is often more technically ambitious than what most consumer-AI companies are doing. If you've been ideologically opposed to defense work, that's your prerogative — but the field is no longer a backwater, and the compute and data budgets are real.
The next visible milestone is the CCA program's first flights of production autonomy stacks on General Atomics and Anduril airframes, expected within the next 18 months. Longer-term, the interesting question isn't whether AI will fly combat aircraft — it will — but whether the Air Force can build the acquisition and certification pipelines to field software at software speed. Right now a new flight-control law takes years to certify. If ACE's model of iterative deployment on VISTA generalizes, that timeline could compress by an order of magnitude, and that's the shift worth watching.
> The kit utilizes a novel interface with the aircraft’s flight controls and mission systems, allowing a pilot to toggle between traditional human control and AI control with the flip of a switch. This ensures a safe, reliable environment for human-on-the-loop experimentation.How, exactly, does t
So, it's a drone with uncessercary baggage such life support and a pilot that can't stand the Gs well.For some reason I am reminded of Dark Star :-)
Without any concrete description of the AI techniques used, one can't help but wonder if they have labeled something like nonlinear model-predictive control as "AI".
The field demo I want to see, Human in the loop to failure switch over: Simulated failure of some form, where (for safety reasons) crew triggers ejection. This defaults the system to autonomous mode and the plane performs a safe landing.
Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.
All Stealth bombers are upgraded with Cyberdyne computers, becoming fully unmanned. Afterwards they fly with a perfect operational record. The Skynet funding bill is passed. The system goes on-line on August 4, 1997. Human decisions are removed from strategic defense. Skynet begins to learn at a geo