An International Publisher for Academic and Scientific Journals
Author Login
Scholars Journal of Engineering and Technology | Volume-14 | Issue-09
From Raw Scientific Data to Trusted Discovery: Agentic AI for FAIR and Reproducible Data Science: A Systematic Review of Data Engineering, Semantic Data Models, Multimodal Analytics, Provenance, and Human Oversight
Jamal Shah, Sunila Kiran, Muhammad Hasnain, Amir Sajjad, Inam Ullah Khan, Ali Hamza, Arsalan Ahmad Khan
Published: Sept. 17, 2026 |
37
11
Pages: 526-550
Downloads
Abstract
Agentic artificial intelligence (AI) is beginning to change scientific data work from a sequence of manually connected tasks into a set of tool-mediated, partially autonomous workflows. That shift creates a practical question: can agents accelerate analysis without weakening the findability, accessibility, interoperability, reusability (FAIR), and reproducibility on which scientific claims depend? We systematically reviewed 218 peer-reviewed papers, preprints, and community standards published from 2016 to September 2026. The evidence was organized into five linked domains: data engineering and AutoML; semantic data models and knowledge graphs; multimodal analytics and retrieval-augmented generation; provenance and reproducibility; and human oversight, trust, and governance. Across these domains, the literature shows clear progress in execution-grounded data preparation, workflow orchestration, semantic retrieval, and multimodal evidence synthesis. It also shows a consistent boundary: systems can generate plausible analyses faster than they can establish that those analyses are complete, reproducible, and scientifically warranted. The most persistent weaknesses are incomplete lineage, unstable dependencies, weak cross-modal semantics, inconsistent evaluation, and poorly calibrated human reliance. We therefore frame trustworthy agentic science as a data-to-decision problem rather than a model-selection problem. Agents should produce claim-level provenance, machine-actionable semantic metadata, executable environments, uncertainty-aware outputs, and explicit escalation points for human review. The review concludes with a staged research agenda and an integrated framework that links FAIR practice, reproducibility engineering, and human governance across the scientific lifecycle.


