The U.S. Army War College attracted considerable attention from military and academic circles in April after publishing the results of an unprecedented experiment involving four prominent artificial intelligence-powered large language models (LLMs): ChatGPT, Gemini, Claude, and Grok. The models were subjected to a comprehensive oral examination similar to the assessment undertaken by the college’s military and civilian students.

All four models successfully passed the examination. However, the results also revealed important limitations in these technologies and offered valuable lessons for military educational institutions worldwide. To understand the significance of this experiment, it is first necessary to examine the nature of education and assessment at U.S. war colleges and comparable institutions.
Military Education at War Colleges: Beyond Lectures and Traditional Examinations
War colleges and national defence colleges represent the highest level of Professional Military Education (PME), for military officers and civilians preparing to assume senior leadership and strategic positions. Their role extends beyond the transfer of knowledge to developing the ability to think critically, analyse complex problems, formulate strategic options, and provide advice to decision-makers.
For this reason, these institutions rely on a combination of academic studies, seminar discussions, individual research, case studies, and strategic exercises rather than traditional examinations alone. Students are expected to connect military history with public policy, national security, and international relations, while demonstrating their ability to think and analyse rather than simply recall information.
The Comprehensive Examination: How Did the Idea Emerge?
For more than a century since the establishment of the U.S. Army War College in 1901, student assessment relied primarily on research papers, academic studies, classroom participation, and seminars. Despite the college’s prestigious standing, there was no final comprehensive examination bringing together the knowledge acquired throughout an entire academic year in a single assessment session.
During the first decade of the 21st century, however, military professional education circles increasingly debated whether traditional assessment methods were capable of measuring the skills required of strategic leaders. A student may be able to produce an outstanding research paper after weeks of work and revision, yet struggle to analyse a complex problem immediately or defend an argument under the pressure of successive questions.
The War College’s adoption of the comprehensive examination came at a time when U.S. forces were engaged in complex operations in Iraq and Afghanistan. This experience reinforced the importance of preparing leaders capable of operating in strategic environments characterised by ambiguity, complexity, and rapid change—skills that are difficult to measure through written research alone. The college subsequently reviewed its assessment system and, in 2013, adopted the Comprehensive Oral Examination as a key requirement for graduation.
For many War College students, the day of the comprehensive examination is regarded as one of the most challenging and stressful points of the programme. Unlike traditional examinations based on written answers or multiple-choice questions, it involves a direct exchange between the student and a panel of faculty members. All students are required to take the examination, and before entering the examination room, they do not know which questions they will face or how the discussion will develop. This requires a comprehensive understanding of the subjects studied throughout the year.
The examination is not intended to measure how much information a student can memorise, but rather how the student thinks. Assessment panels seek to examine several core abilities, including analysing complex strategic problems, connecting military history with contemporary issues, evaluating different options and discussing their advantages and risks, constructing evidence-based arguments, defending intellectual positions under pressure, and revising assessments when new information emerges.
In this respect, the examination resembles the situations a commander or strategic adviser may encounter in practice, when required to provide a rapid professional assessment to a senior commander or political decision-maker.
AI Enters the Classroom
In recent years, the U.S. Army War College has not been isolated from the transformation brought about by generative artificial intelligence. Like civilian universities, military educational institutions have begun examining how these tools can be used in education, research, and analysis. Particular attention has focused on the ability of large language models to summarise information, analyse texts, propose alternatives, and assist in preparing research and academic papers.
The key question, however, remained: if these models are capable of producing convincing texts and sophisticated analyses, could they pass the same examinations administered to War College officers?
This question led to the experiment conducted by several members of the college’s faculty during 2026.
An Unprecedented Experiment
As noted earlier, researchers subjected four of the most prominent commercially available AI models to an examination designed to simulate the comprehensive oral examination administered to students at the U.S. Army War College. The researchers used a benchmark known as MilBench, an assessment framework designed to measure AI performance in areas considered fundamental to strategic military education, including critical thinking, strategic analysis, the use of military history, connecting different concepts, and defending arguments during discussion.
The models underwent extended questioning sessions similar to those faced by actual students. The assessment did not consist merely of separate questions; instead, it involved continuous dialogue requiring the models to maintain intellectual coherence and provide consistent, substantive, and in-depth responses throughout the discussion.
When faculty members decided to subject the AI models to the examination, they were not testing their ability to recall information. Rather, they were assessing whether the models could perform the same task expected of a strategic leader: analyse a complex problem, connect diverse pieces of information, and defend their conclusions during an extended and evolving discussion.
For this reason, the results of the experiment were particularly significant. The measure of success was not knowledge alone, but the ability to exercise strategic thinking as defined by one of the world’s most established military educational institutions.
Results: Impressive Performance and Clear Limitations
The results showed that all four models successfully passed the examination, which researchers viewed as an important indication of the significant advances made in artificial intelligence applications in recent years. However, the results also revealed clear differences among the models, with one outperforming the others notably in analytical quality and its ability to address complex questions. Claude received an A+, while the other models received B+ grades. The most significant observation, however, was that the performance of all four models began to decline as the questioning sessions became longer. As the dialogue continued, their responses became more repetitive, less substantive, and less capable of developing ideas or constructing new arguments.
Researchers attributed this phenomenon to the current computational and technical limitations of large language models. Some researchers and analysts have described it as a form of “Artificial Cognitive Fatigue,” in which performance quality declines as dialogue and questioning become prolonged.
Perhaps the most important finding from a military perspective is that commanders are not normally assessed on their ability to provide a good answer to a single question. Rather, they must maintain the quality of their thinking and the consistency of their judgments throughout lengthy discussions, assessments, and decision-making processes.
This highlights the distinction between the ability to produce a convincing response and the ability to exercise sustained professional judgment in environments characterised by pressure and uncertainty. In other words, the models demonstrated a strong ability to retrieve and analyse knowledge, but faced greater difficulty maintaining the same level of performance throughout a lengthy and complex dialogue.
What Do These Findings Mean?
The researchers concluded that artificial intelligence can serve as a highly valuable support tool in advanced military education, but it does not constitute a substitute for the student, faculty member, or military commander.
Current models can provide information and preliminary analysis at very high speed and may assist in exploring alternatives, reviewing ideas, or summarising studies. However, they still lack certain characteristics associated with human judgment, including an understanding of the full context, assessment of political risks, appreciation of human factors, and decision-making under pressure.
Accordingly, the researchers recommended treating artificial intelligence much like a capable staff officer—one that nevertheless requires direction, supervision, and verification of its outputs.
Lessons for Educational Institutions
The experiment offers several important lessons for military and security educational institutions. Foremost among them is the recognition that artificial intelligence has become an educational and professional tool that should be understood and integrated carefully into curricula and educational programmes.
At the same time, the importance of critical-thinking and strategic-analysis skills has become greater than ever. If machines can retrieve information rapidly, the real value of a leader lies in the ability to interpret that information and apply it effectively to decision-making. The experience also highlights the importance of oral examinations and interactive discussions as methods of assessing students’ actual capabilities. Such methods test thinking in real time rather than merely the ability to prepare written responses.
The experiment also raises an important question about the suitability of assessment methods used by some advanced military educational institutions. As AI tools become increasingly capable of assisting with research papers and reports, greater emphasis will need to be placed on assessment methods that measure direct thinking, oral analysis, and the ability to defend positions in real time.
Finally, the experiment highlights the need for clear frameworks and policies governing the use of artificial intelligence in military educational institutions. Such frameworks should enable institutions to benefit from AI capabilities while preserving academic integrity and developing students’ independent-thinking skills.
Conclusion
The U.S. Army War College experiment demonstrated that artificial intelligence applications are now capable of passing examinations that, until recently, were considered the exclusive domain of humans. At the same time, it revealed clear limitations that continue to separate information processing from comprehensive strategic judgment.
The true value of the experiment may therefore lie not in demonstrating that AI can pass a high-level academic examination, but in reminding us that preparing leaders in the 21st century is no longer primarily about possessing knowledge; it is about using that knowledge wisely.
While information has become accessible to everyone, sound professional judgment and the ability to make decisions under conditions of uncertainty and pressure remain among the human qualities that AI has yet to fully replicate.
By: Major General (Ret.) Khaled Ali Al-Sumaiti




