Agentic AI in OSINT: Taxonomy & Reliability Gaps
Abstract
The rapid growth of publicly available digital information has rendered manual open-source intelligence (OSINT) analysis insufficient for contemporary intelligence, cybersecurity, and Cyber Investigation requirements. Large language models (LLMs) and agentic AI systems, which select tools, perform multistep reasoning, and iteratively produce intelligence, have emerged as promising responses, yet published capability demonstrations have substantially outpaced the evaluation infrastructure required to validate operational deployment. This survey systematically reviews 74 unique studies on the application of agentic AI, generative AI, and LLMs to OSINT, cyber threat intelligence (CTI), and Cyber Investigation. Its contribution is fourfold. First, it treats agentic AI as a distinct analytical category rather than a variant of LLM prompting, organising the literature through an 11-category taxonomy spanning LLM foundations, agentic architectures, retrieval-augmented generation (RAG), knowledge graphs, prompt engineering, domain adaptation, evaluation benchmarks, and risk. Second, it establishes the hallucination validation gap as a corpus-level finding: although hallucination is named a reliability concern in more than twenty studies, end-to-end hallucination is empirically measured in only one OSINT-specific system, a RAG-augmented architecture reporting a 4% rate under favourable, non-reproducible conditions; the reasoning-error and factual-correction results reported elsewhere are measured in general-domain question answering, not on OSINT hallucination, and do not close this gap. Third, it maps the corpus onto the OSINT workflow lifecycle, showing that collection and analysis are well served while verification, reporting, dissemination, and decision support remain systematically underexplored. Fourth, it derives a ten-point research agenda, covering evaluation, benchmarking, hallucination measurement, adversarial robustness, dark-web coverage, multimodal processing, and governance, directly from the gaps the corpus exposes. The review further finds that no standardised, open, communityadopted benchmark exists for cross-study comparison of OSINT AI systems, and that agentic systems are evaluated exclusively under benign conditions despite documented adversarial threats. It concludes that a structured humanAI co-pilot model, in which LLMs support collection and triage while analysts retain responsibility for verification, reporting, and decision support, is the most defensible near-term deployment architecture under the current evidence base.