Date of Award
Spring 2026
Document Type
Honors
Degree Name
General Beadle Honors Program
First Advisor
Tyler Flaagan
Abstract
This study evaluates the ability of locally hosted large language models (LLMs) to autonomously preform multi-step offensive tasks against a vulnerable target. Tested in this study are three publicly obtainable LLMs; Qwen3.5, Llama3.1, and Gemma4 configured to use Model Context Protocol (MCP) in a host-only virtual network comprising of a Kali Linux attacker system and a Metasploitable 2 target. Each model executed five independent attempts across two phases; reconnaissance (service and port discovery) and exploitation (attempt to open a session). This study measured phase success rate, tool call counts, redundant and erroneous calls, and human interventions. Results indicate that all models can complete reconnaissance reliably; Gemma4 demonstrated the highest practical autonomy and efficiency in exploitation, Qwen3.5 succeeded but exhibited frequent redundant and corrective tool calls, and Llama3.1 demonstrated understanding of the necessary steps but lacked actual execution of steps with increased task complexity. This study assists cybersecurity and AI researchers in understanding the capabilities and risks of locally hostable AI models, how threat actors using locally hosted LLMs can utilize similar capabilities to private commercial AI models, and the implications this has for defenders.
Recommended Citation
Ahrendt, Bryston, "Measuring Atonomy: Local LLMs Orchestrating Reconnaissance and Exploitation" (2026). Honors. 21.
https://scholar.dsu.edu/honors/21
