lakeFS and E2B present a receipts-processing pipeline designed to let an AI agent work on real data without directly modifying production state. E2B runs the agent inside a Firecracker microVM with a Linux kernel, filesystem access, external API access, and support for generated code. lakeFS creates a zero-copy branch for each run and mounts it inside the sandbox through lakeFS Mount and FUSE, so the agent uses ordinary file operations without importing a storage client or lakeFS API. The demonstration starts with a deliberately messy inbox containing receipt and invoice images in multiple formats, duplicates, a corrupt file, a non-receipt image, and an unsupported text file. The agent processes the data in triage, extraction, and validation phases, committing the mounted filesystem after each phase to create an auditable history. It uses a vision model during extraction and can generate validation logic from a plain-English specification at runtime. When currency or totals are uncertain, the agent flags rows for human review; the host pauses the sandbox, resumes it after a decision, and commits the result. A server-side lakeFS pre-merge hook independently checks the committed ledger, including currency, policy caps, line-item totals, dates, and invoice-number uniqueness, and only valid output can merge into main. The demo reports that a prebaked E2B template reduced startup from about 19 seconds to about 1 second, while lazy fetching accessed only a small portion of a 4.88 GB, 5,000-file dataset. The authors also tested a 30-minute pause and resume, finding that the mount remained usable and could read cached and new files and commit changes. The open demo requires an E2B account and a lakeFS instance with Mount enabled; the article says lakeFS Mount is available on lakeFS Enterprise.
