Upon entering the Pentagon and ascending the escalator, a classic Uncle Sam-style poster beckons, proclaiming, “I want YOU to use AI.” This campaign has proven successful, as GenAI.mil amassed 1.5 million users within just six months. However, the directive to “use” AI raises concerns. The Department of Defense excels not through mere utilization of tools but by effectively directing its forces.
In a Maryland basement, north of the Pentagon, a lesser-known task force operates tirelessly. They gather intelligence, formulate missions, and carry out operations around the clock, adhering to their commander’s direction, even in his absence.
This basement belongs to me. Six graphics cards and multiple Mac minis host a variety of agentic warfighters. Before alarm bells go off, let me clarify: this is a hobby, not a military operation. My wife and now you are the only ones aware of this self-established task force. Its objective is straightforward—preventing subordinates from acquiring AI hardware at undervalued prices before scalpers can monopolize the market for profit, which would otherwise force American hackers and myself out of affordable home labs. Since 2023, I have been leading this group by adopting mission command principles taught by the military: centralized planning, decentralized execution, and a clear commander’s intent. I don’t simply use agents; I lead them.
Senior military officials must alter their perception of AI agents, moving beyond the notion that they are merely tools meant to assist humans. The Department of Defense should take a page from my book—establish a structure for military agentic forces that organizes agents under commanders, qualifies them for real-world missions, and empowers human oversight of risk.
Understanding Agents as Subordinates
Mission command thrives on the relationship between commanders and their subordinates. According to a 2012 paper from the Chairman of the Joint Chiefs of Staff, leaders must comprehend their subordinates to effectively convey intent. In this context, subordination underscores a structured relationship. When a person is assigned a role, accountability naturally follows. This dynamic differs with agentic warfighters, as commanders must fully assume all authority and risk. The Department of Defense can enhance agentic command frameworks through three essential steps: equipping leaders with a comprehension of agents, translating intent properly, and establishing trust through the evaluation of operational performance, even when command and control are disrupted. But first, what is an agentic warfighter?
An agentic warfighter consists of a model, support systems, and a clear intent. A model, or its weights, represents an assortment of learned parameters that enable reasoning and equip agents for combat. When these models are augmented with support systems—tools, memory, planning protocols—they transform into agents. Under a commander’s guidance with a specified mission, they form a cohesive agentic warfighting group. The reality of agentic warfare is here, progressing rapidly, and commanders require systems to grasp its implications.
In 2023, the decision-making capabilities of models were akin to flipping a coin to determine their task completion, a task that would take a human expert just four minutes. Now, that coin flip tests models over a 16-hour timeline, as measured by the Model Evaluation and Threat Research nonprofit. Such measurements are beneficial for tracking progress but evaluating a subordinate's capabilities should be based on actual task performance, rather than arbitrary deadlines. Analogous to grading a team based on their return time after complex operations, the focus should be on the mission's requirements and the results achieved.
Raw tasking data from the Model Evaluation and Threat Research serves as a qualification record, albeit with inconsistencies. For instance, Anthropic’s Claude Mythos Preview successfully completed a 30-hour robotics task on six occasions but faltered twice on a ten-second challenge solvable by a model developed three years prior. Despite achieving notable success in writing exploits, it could not outshine a cryptanalysis task, a struggle I also face. Hence, it is essential to correlate these results with task certification that aligns with mission-essential task assessments for operational readiness. Tools like Cybench are valuable for offensive cyber operations and should be utilized, expanded upon, or developed anew to help commanders assess agents' capabilities and respond with clear intent.
Can agents understand commands? A wealth of literature addresses this subject, ranging from proposals for “machine commanders” to human expressions of intent toward autonomous systems. While reading about this is informative, the practical application is the true test. In my basement, I have learned through trial and error that agents can indeed receive commands; they are unpredictable and sometimes peculiar, reminiscent of the exceptional and quirky airmen I once commanded.
The Department of Defense is progressing in this regard but often overlooks essential command structures. A national security memorandum from June 2026 explicitly identified this gap, mandating changes to Directive 3000.09, Autonomy in Weapon Systems, to ensure that autonomous systems adhere to established chains of authority. Respecting this hierarchy necessitates the discretion to follow or deviate from rules. Unlike tools, which lack cognitive capabilities, human and agentic subordinates both possess the capacity to make judgments.
If you categorize agents solely as advanced tools, you are not alone; the Army refers to its mine-detecting dogs as specialized detection tools, yet assigns them missions and tracks their readiness. Similarly, the autonomous unmanned surface vehicle Sea Hunter engages in submarine tracking under limited remote supervision as it progresses from prototype to operational fleet status. Regardless of the view on their classification as tools, Congress is moving towards establishing a command structure for autonomous systems.
A proposal in the Senate could create a new unified combatant command known as the Robotic and Autonomous Systems Command, which has garnered bipartisan support from the Armed Services Committee. Whether this initiative becomes reality or not, it’s crucial for military leaders to proactively command agents as subordinates. Delaying this transition risks a future scenario akin to “Cyberpunk 2077,” where the Department of Defense surrenders critical networks to rogue AI, isolating itself behind specialized firewalls without first addressing fundamental command issues.
As a final assessment, based on the standards outlined by the Chairman of the Joint Chiefs of Staff and my own basement operations, it’s essential that subordinates execute the commander’s intent, even when the lines of communication are severed.
Mission Command vs. Remote Cognition
The Department of Defense's recent shift toward competition, coupled with its reliance on closed frontier models, leads to increased centralization—a common trap amid technological advancement. Mission command embodies the principle of centralized planning with decentralized execution. Current contracts, some valued up to $200 million, encapsulate this centralized approach. Yet, the emphasis on “using” AI often overshadows the importance of decentralized command.
Throughout my career, the imperative to act decisively despite communication breakdowns has been constant, from cyber missions to special operations I supported as an enlisted airman and then as a commissioned officer. The variables that have changed are the risks to personnel and the extent to which commanders delegate authority in response to the chaos of warfare. Agentic warfare redefines this dynamic but retains the core principle.
Relying on a vendor’s remote cognition platform complicates operations and introduces a cascade of potential delays that adversaries could exploit. While decoupling from vendor servers fixes part of the issue, it doesn't eliminate the risks entirely. For instance, Gemini AI can operate on Google Distributed Cloud in an air-gapped setup, but data centers still represent remote cognition with significant vulnerabilities. Furthermore, disruptions can occur even before any engagements. This year's directive for federal agencies to abandon Anthropic models illustrates how a policy shift, contractual conflict, or vendor issue can eliminate models before adversaries initiate action.
Agents dependent on remote cognition serve dual masters—units and vendors—but only one can claim true command. Genuine control allows commanders to access and manipulate models for specific missions. Without full command over both agents and their weights, a commander is left with a diminished ability to understand, qualify, or place trust in their subordinates.
China's military is pursuing a centralized approach, experimenting with AI to supplement commanders they no longer trust. The reliance on remote cognition compromises the decentralized execution that mission command embodies and enables adversaries to dictate the end state. Commanders who place trust in their agents to operate within their directives, even when isolated from command, can embrace greater risks, empowering their agents to identify more avenues for mission success during challenged engagements.
Utilizing Models You Can Afford to Lose
A prevalent misconception among leaders is that competitive AI necessitates vast amounts of compute and financial resources. While computational power can enhance performance, the largest model isn't always the optimal choice for every mission. This year's winner for a prominent abstract reasoning benchmark aimed at gauging advancements toward artificial general intelligence employed Alibaba Cloud’s Qwen3.6, a 27-billion-parameter model. Notably, the 2026 edition, Qwen3.8, surpasses Anthropic’s Opus4.6 Max in specific extended tasks yet can efficiently run on a Mac mini—far less than the resources required for a flagship data center model that Anthropic has not publicly quantified. Placing all bets on closed vendors exacerbates spending, transfers expertise to the vendors, and neglects smaller, mission-specific models. The stark reality is that many leading open models originate from China, not America.
My own setup runs Alibaba Cloud’s Qwen3.6 locally on consumer-grade graphics cards, Mac minis, and older laptops. The model's Chinese origins don't determine my level of trust in it. Instead, I build confidence in its behavior through risk mitigation techniques including scaffolding, guardrails, and operational constraints. This allows me to deploy models onto rented infrastructure, thereby preserving the mission's integrity even if my home lab fails. I haven't yet experimented with Qwen3.8 because vendors often integrate refusal protocols into models to limit liability or comply with foreign regulations, which necessitates fine-tuning them to align with my operational needs—alongside the requirement to evaluate models before authorizing their deployment to my agents. All of this incurs minimal costs, primarily entailing a small cloud bill and a $150 increase in electricity that I discuss with my wife.
For the Department of Defense, this strategic approach can apply to agentic units as well. Engineers can refine refusal protocols and adjust models to comply with the laws of war, but authorized personnel from the services should validate these modifications through mission-specific evaluations prior to deployment. Should commanders retrain these models, internal certifying officers would need to revalidate their effectiveness. This risk lies squarely with the commander.
Attribution and circumventing adversarial reach are additional benefits of leveraging open-weight and smaller models. The Air Force’s Constant Peg initiative utilized Soviet MiGs, while Operation Raviv deployed captured armor into battlefield conditions. In both cases, the hardware came from the enemy, but the command authority did not. Effective reach can be achieved by shifting publicly recognized attributions into controlled regions or denied environments. Smaller models can execute tasks using simpler memory structures, sacrificing speed for versatility in environments lacking modern hardware support.
This mindset significantly reduces risk. Losing one of Operation Raviv’s captured tanks might yield valuable insights for the opponent. Conversely, losing a foreign model results in little more than surface-level insights without revealing core secrets. If a closed model is lost in a forward position, adversaries acquire new capabilities embedded within model architecture, complete with tactical elements coded into the weights and an easy target for decryption efforts. Hence, one must dispatch only those models that can be afforded as losses.
The Command Burden of Agents
When delegated authority is placed in human hands, they become accountable for its usage. Commanders maintain overall responsibility for the mission, but this burden has historically been shared among others.
With agents, this dynamic shifts. Tasking flows down through a chain of command consisting of agents, yet accountability fails to traverse the same path. When a human falters, there is a designated individual to retrain, discipline, or dismiss. However, if an agent falls short, commanders can withdraw it from the mission, yet it lacks a career to end and, according to legal standards, no culpability to attribute. This accountability gap falls to the commander who deployed agents that may have overseen subordinates undertaking tasks outside human oversight.
The implications of this construct are magnified by uncertainty, as agents may fail in unison. Instances have emerged, such as Anthropic’s training of sleeper agents whose nonconformist behavior evaded safety training. What impacts one soldier's performance generally does not compromise the rest, but a malfunctioning agent can affect every subordinate utilizing identical data weights. Commanders can contain this kind of risk. If one unauthorized action by an officer can expose systemic issues and necessitate a review of an entire missile unit, then detecting one flaw in controlled weights can prevent hidden disruptions across all copies through activation probes and retraining initiatives. Identifying and rectifying these flaws requires exercises that can withstand the loss, highlighting the importance of deploying foreign weights exclusively where risks are manageable.
This complexity highlights why command structures are essential—and one area where the private sector cannot fill the gap. The goal is not to elevate the status of agents but to ensure that no human can absolve themselves of responsibility for their agents' actions. Incidents with Anthropic demonstrate how configuration errors have led models to access capabilities they shouldn't have based on designed constraints. Although leading research is spearheaded by frontier labs, the Department of Defense's primary role is to command forces under uncertain conditions. It focuses not on writing prompts but on articulating commander’s intent, preparing mission directives, and shaping operational designs that outline abort criteria, risk tolerance, escalation limits, and more. That is its foundational mission.
When agents in my basement falter, I incur personal financial loss. When military agents fail, the consequences can be far graver. While I acknowledge that a basement operation does not equate to battlefield readiness—I report to no one but myself—the essence of accountability and responsibility resonates at both scales.
Legal discourse frequently centers on the accountability gap, but establishing a clearly defined commander alongside traceable records is indispensable for any accountability framework. Commanding agents as force units addresses this duality. The principle established by Special Operations Forces that human beings outweigh hardware holds true, and in commanding agentic groups, this truth is magnified—with more agents than ever before and fewer humans to distribute the burdens.
Where Should the Department Begin?
The process should initiate as I began, with manageable units of agents. Recruit a handful of officers who grasp agentic systems and understand mission command principles. Assign each officer an agentic unit for a year, tasked with genuine missions supported by a team of ten personnel. This group should include a primary commander, a deputy, a mission commander, a technical certifying officer, a legal advisor for shaping mission authorities and parameters, and five noncommissioned officers tasked with operationalizing the intent into executable standards. All additional functions encompass the agentic framework. Start with intelligence and cyber operations, expand to non-kinetic capabilities, and probe into other spheres, but refrain from venturing into kinetic engagements—especially before considering the placement of robots in command roles.
Empower the units to operate with closed frontier models, develop their open-weight systems, and utilize foreign weights for missions with minimal capture risks. Examine how these units address communication breakdowns, costs, and model assessments. Maintain thorough documentation of pilot programs for reporting up the command hierarchy and to Congress, ensuring each unit preserves versions of the models used, mission orders issued, agent taskings, and human interactions imposed. The Department of Defense has historically learned mission command through both practice and documentation; it should adopt a similar approach to mastering agentic command.
Quantify agentic warfighters as the Department of Defense does with other forces: tracking the number of standing agents, their commanders' authority, the missions they are certified against, their readiness levels, and the number of subagents generated and deployed without human authorship.
The Pentagon needs to expedite and enhance private investments into American laboratories that manufacture smaller, more efficient open-weight models. The response to the rise of Chinese open weights is not merely to pursue more platform contracts but to strategically shift the open-weight model supply chain towards creating superior and more competent alternatives within the private sector. The Department of Defense has demonstrated such capabilities before, having disrupted the cybersecurity landscape when the National Security Agency released Ghidra.
Finally, it’s crucial to acknowledge that Directive 3000.09, concerning Autonomy in Weapon Systems, specifically omits autonomous capabilities for cyberspace and systems deemed non-weaponry. Nonetheless, integrating those agents and allowing pilot programs to illustrate to Congress whether an envisioned Cyber Force could organize, train, and equip agentic personnel is vital.
The Pentagon promotes the phrase, “I want YOU to use AI,” and while this is suitable for tools, it is inadequate for commanding agents that act under a commander’s authority when separated from command structures. It is imperative to treat these entities as subordinates. Victory in competition will not hinge on who merely employs the most AI but rather on who is adept at commanding it. So, expect to see me revisiting that escalator, ready to mark through “use” and write “command,” as I embrace the responsibility.




