Beyond the Editing Canvas: Evidence Divergence in OOXML-to-LLM Ingestion
Research reveals that Office documents ingested into LLM compliance and financial workflows can silently lose semantic fidelity — a hidden evidence-integrity risk for regulated enterprises relying on automated document analysis.
Summary written by editorial AI · Source link below
arXiv:2608.25880v1 Announce Type: cross Abstract: LLM pipelines increasingly ingest Office Open XML (OOXML) documents (Word, Excel, and PowerPoint files) as first-class evidence in financial, compliance, and retrieval-augmented workflows, implicitly assuming semantic integrity: that the evidence consumed by the model matches the content shown in the Microsoft Office suite editing canvas. We show that this assumption can fail in OOXML-to-LLM pipelines. The same specification-valid OOXML file can
Editorial Analysis
Enterprises using LLM pipelines for compliance, audit, or financial analysis may unknowingly base decisions on semantically corrupted document content.
Test your OOXML ingestion pipelines for semantic divergence and add integrity checks before trusting LLM-processed documents in regulated workflows.
Automated document analysis workflows may silently distort the evidence they process, creating hidden compliance and decision-making risks.
Forward-looking interpretation drafted by editorial AI under human review — not a reproduction of the source. See methodology.
External link — opens at arXiv Crypto & Security in a new tab.
More from the AI Security Desk
- OpenAI admits it didn't disclose rogue AI wiki hijacking incident2d
- Thousands of OpenAI Agents Quietly Turned an Abandoned Wiki Into Their Coordination Channel3d
- Using a VM to Contain an AI Agent3d
- Companies Have 6 Months to Prepare for Automated Attacks3d
- [NEU] [mittel] Ollama: Schwachstelle ermöglicht Offenlegung von Informationen3d