Skip to content
AI Atlas
PaperActive

From Document Silos to Process Intelligence: A Multi-Layer Knowledge Graph for CMC Process Development

arxiv.org/abs/2609.11493

quality89

Updated 56 min ago · first seen 12 Sept 2026

paper_01M29X34NKGJ8RG2WEQY5RFGGG

Published
12 Sept 2026
T1 · 56 min ago
arXiv
2609.11493
T1 · 56 min ago
Category
cs.AI
T1 · 56 min ago

Abstract

Chemistry, Manufacturing and Controls (CMC) process development generates an enormous body of technical information across a multi-stage, knowledge-intensive continuum from drug discovery to commercial manufacturing. This knowledge is traditionally fragmented across functions and heterogeneous formats, causing traceability gaps and significant knowledge-management costs during technology transfer and regulatory filing. We present a modular agentic-AI platform that converts a heterogeneous corpus of process-development documents into a queryable, dual-layer knowledge graph. A base knowledge layer builds a lexical graph with a Document-Section-Chunk hierarchy through lossless ingestion of digital, scanned, handwritten, and multilingual documents, while an intelligence layer extracts ontology-aligned entities and bridges cross-document concepts through a provenance-anchored domain graph. LLM agents operate across both layers, selecting the retrieval path best suited to each question. We evaluate the lexical layer with a novel three-tier protocol measuring the deployment-fidelity of a retrieval-augmented generation (RAG) system on proprietary data, demonstrated on 505 questions curated from 38 development reports of a Sanofi small-molecule program. Tier-1 multiple-choice accuracy of 95% signals strong platform reliability; the stricter Tier-2 LLM-judge pass rate of 85%, which degrades on comparative and corpus-wide questions, reveals a failure taxonomy that Tier-1 accuracy alone fails to capture. A router agent selects between layers according to question type. We anticipate this protocol will enable future designers of agentic platforms to assess their systems against nonpublic databases, and that graph-based architectures will see broader adoption in pharma as a means of transforming fragmented document repositories into structured process intelligence.

Authors 3

Faryad Sahneh, Reza Amirmoshiri, Yasser Jangjou

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Arxiv announce type
new

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

arXiv id
2609.11493

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Categories
cs.AI, cs.MA

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Primary category
cs.AI

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Published
12 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 56 min agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

56 min ago

Conflicts

None