#UAE #physicalAI - United Arab Emirates University, in collaboration with the Digital Future Institute of Khalifa University and Abu Dhabi’s Technology Innovation Institute (TII), have released PhysAI-Bench, a new benchmark testing whether AI foundation models can make sound real-time decisions as autonomous drone pilots. Leveraging 10,178 decision points drawn from real UAV (unmanned aerial vehicle) mission traces, the benchmark was used to test 29 models from 14 AI organisations, including OpenAI, Anthropic, Google, Meta and xAI. The best-performing model, OpenAI’s GPT-5.3 Chat, correctly selected the right action just 52 percent of the time on unseen test data, with every model scoring lower on fresh questions than on those used to tune it.
SO WHAT? - Overall frontier models performed poorly in the PhysAI-Bench tests. Even the top score of 52 percent among the models tested is inadequate for mid-mission autonomous decision-making during drone operations. The bigger finding is perhaps the gap between how models perform on the small set researchers used to tune them, versus fresh, unseen scenarios (a gap that showed up in all models tested). The new benchmark and research paper provides food for thought for anyone planning to deploy AI models in physical, safety-critical settings.
KEY POINTS:
UAE University, in collaboration with Khalifa University and Abu Dhabi’s Technology Innovation Institute (TII), has released PhysAI-Bench, a new benchmark that tests real-time decision-making of AI models when they are being used as autonomous drone pilots.
PhysAI-Bench evaluates 29 foundation models from 14 AI organisations on their ability to select the correct next action for an autonomous drone, using 10,178 decision instances built from real UAV mission traces.
GPT-5.3 Chat achieved the highest accuracy in the benchmark tests at 52.00%, followed by GPT-5.2 Chat at 49.40% and xAI’s Grok 4.5 at 49.07%. The models were scored on a fixed set of 500 previously unseen test questions.
Every one of the 29 AI models performed worse on the held-out test set than on the smaller development set used to tune its settings, with gaps ranging from roughly 2% to 26%.
Model size alone didn’t predict performance. Several smaller, open-weight models such as Mistral Small 2603 and the Nous Hermes family managed to match, or beat, much larger frontier models while running faster and more efficiently.
Decoding temperature, a setting that controls how random a model’s answers are, made little difference to the strongest models. However, the number of examples given beforehand (few-shot prompting) helped most models, though by widely varying amounts.
The NVIDIA Nemotron family scored particularly poorly, largely because their answers failed to follow the required format in a large share of test cases, with parse success rates as low as 14%.
Each decision instance captured mission objectives, physical constraints, sensor readings, tool calls and simulated 6G network conditions like latency and packet loss, designed to reflect what a real onboard AI agent would actually see.
The research team plans to extend the approach beyond drones to robotics, self-driving vehicles and industrial automation.
Principal researcher on this project were Mohamed Amine Ferrag PhD, Senior Member, IEEE and Associate Professor at UAE University; Abderrahmane Lakas, Senior Member, IEEE, professor and Assistant Dean for Research & Graduate Studies at UAE University; Merouane Debbah, Fellow, IEEE, Professor at Khalifa University and Founder of the university’s 6G Research Centre; Manu Perumkunnil, Interuniversity Microelectronics Centre (IMEC), Belgium; and Norbert Tihanyi, PhD, Technology Innovation Institute (TII).
ZOOM OUT - The PhysAI-Bench research builds on earlier work from UAE University and Khalifa University regarding the use of frontier models for UAV autonomous decision-making. In November 2025, the two institutions released UAVBench, an open dataset of 50,000 validated drone flight scenarios. Researchers paired flight scenarios with 50,000 multiple-choice reasoning questions spanning ten domains, from aerodynamics to ethical decision-making. The UAVBench project tested 32 models, including GPT-5 and Gemini 2.5 Flash, and found strong perception and policy reasoning but persistent weaknesses in resource-constrained and ethics-aware decisions.
[Written and edited with the assistance of AI]
Source: UAE University, Khalifa University, TII
LINKS
PhysAI-Bench research paper (arXiv)
PhysAI-Bench benchmark schema (Github)
Read more drone news:
Researchers release 50,000 drone scenario LLM benchmark (Middle East AI News)
Smart Municipality Eye initiative deploys UAVs across parks (Middle East AI News)
e& UAE deploys autonomous drones for base stations (Middle East AI News)
Keeta Drone secures UAE’s first BVLOS licence (Middle East AI News)
Abu Dhabi launches AI & robotics climate tech venture (Middle East AI News)



