Full-text search across every page
Search the content of every processed document, not just filenames. Image-only PDFs and photographed pages become searchable text through OCR.
Formalbyte e-Document Manager turns scanned paper into a full-text searchable electronic archive. Built on open source, classified by OCRskill agentic OCR, deployed on a virtual machine your company owns.
One time implementation fee, then 25 EUR per month. See what is included
3 results in 0.12 s
Invoice 4127 — Acme Distribution
Service estimate 8842 — Acme Distribution
Delivery note 1190 — Acme Distribution
1 document is waiting for human review. Nothing uncertain is filed on its own.
up to 95%
of documents filed automatically, without human input
seconds
to find any document by the words printed on it
zero
documents stored outside your own infrastructure
1 week
from kickoff call to a working archive, typically
Why now
Cardboard boxes labelled 2019, Invoices or, worse, Misc. One document lost during an audit, a warranty dispute or a GDPR request can cost more than the entire system below.
Capabilities
Search the content of every processed document, not just filenames. Image-only PDFs and photographed pages become searchable text through OCR.
OCRskill reads each document and fills in document type, issue date and correspondent, up to 95% of the time without anyone typing a label.
Unreadable, unusual or out-of-scope documents are flagged for a person instead of being filed blindly into the archive.
Files land in a consistent year, document type, month and correspondent path directly on your own storage.
Diacritics and mixed-language documents are handled in the same pipeline, including scans made from imperfect originals.
Individual accounts, permissions per document type and a complete history of who changed what and when.
Staff scan whole stacks into a shared SMB folder. The system collects the files and processes them without further clicks.
Search and manage the archive from a normal web address over HTTPS. There is nothing to install on every laptop.
How it works
Staff feed documents into the office scanner, stacks at a time, and save them to the shared scans folder.
The system picks up new files automatically and extracts the text, including from image-only PDFs.
OCRskill determines the document type, the real issue date printed on the page and the correspondent.
Anything unclear or unrelated to your archive is flagged for review rather than filed incorrectly.
It is filed into the right folder and is instantly searchable for every authorized user in the browser.
Where the file lands
/archive/2026/Invoice/07-July/Acme-Distribution-Invoice-4127.pdf
Year, document type, month, correspondent. If you know what the document is, you already know where it lives. No more Misc folder, no more scanned_document_0002.pdf.
Data sovereignty
Confidentiality is the foundation of the architecture, not a setting you switch on later. The AI reads text through an encrypted tunnel and forgets it immediately.
A maintainable Docker Compose deployment on an Ubuntu VM you provide, wired into the network folders your team already uses. No exotic dependencies, no vendor lock-in.
AI with a human checkpoint
Automation absorbs the repetitive metadata work. People keep the final say on uncertain or exceptional documents, so a wrong guess never becomes a wrong archive.
Two editions
On an archive of 40.000 documents, the gap between 55% and up to 95% automatic classification is measured in thousands of hours of manual work.
Rules only, no AI module
OCRskill agentic classification
Investment
Initial Setup
One time cost
Includes:
Maintenance and AI processing
25EUR / month
Indicative pricing, excluding VAT. Importing a large historical archive is quoted per processed page after a short scoping call, and you only pay for the pages actually processed.
What you get back
20 min to 5 sec
One document lookup
A trip to the archive room becomes a search box. Multiply by every lookup your team makes in a month.
~1.900 hours
Manual filing avoided
Tagging 40.000 documents by hand costs roughly 2.000 hours. At up to 95% automatic classification, almost all of it disappears.
30+ boxes
Shelving reclaimed
80.000 scanned pages is more than thirty archive boxes you no longer need to store, dust or move.
days to minutes
Audit preparation
Instead of pulling binders for a week, export the requested date range and hand it over.
Illustrative figures for an archive of roughly 40.000 documents at about two pages each. Your own numbers depend on volume, scan quality and how often documents are retrieved.
A strong fit for
Invoices, receipts, purchase orders and bank documents, ready for any inspection.
Contracts, case files and client correspondence, searchable by any clause.
Invoices, service estimates, warranty files and supplier correspondence.
Employee files and administrative paperwork with permissions per document type.
Delivery notes, CMRs and transport paperwork found by shipment reference.
Clean, date-stamped archives and fast answers to GDPR or regulatory requests.
Rollout
Thirty minutes to map document categories, users, storage and the current scanning flow.
The stack goes live on your virtual machine, usually in under a week, storage and VPN included.
Start with one category and one week of new documents, then confirm classification and folder paths.
Import the historical archive in controlled batches, train the team, retire the filing cabinets.
FAQ
It is a self-hosted electronic document archive built on open source and extended with OCRskill agentic OCR. It runs on an Ubuntu virtual machine on your own infrastructure, indexes every scanned page for full-text search, and fills in the document type, issue date and correspondent automatically so nobody has to type labels by hand.
On your own storage. The platform runs as a virtual machine on your infrastructure, and both the scan intake folder and the finished archive live on company storage shared over SMB. Formalbyte does not host your documents.
No. The AI module uses a private large language model hosted in the European Union. Document text travels encrypted through a VPN tunnel, is classified, and the result returns to your system. Documents are never stored on the AI server and never reach a cloud provider outside the EU. A full GDPR storage and processing agreement is included.
With the AI module we typically see up to 95% of documents classified correctly without human input, depending strongly on scan quality, language and how consistent your document types are. A rule-based setup without AI lands closer to 55%. Anything the system is unsure about is flagged for human review instead of being filed blindly.
Usually under a week after a 30-minute scoping call: installing the virtual machine, connecting SMB storage, configuring the VPN, defining your document types and training the team. The slow part is never the software, it is scanning the historical boxes.
Yes, as long as the scanner can save PDFs or images to a network folder. Staff feed whole stacks into a document scanner, the files land in the shared scans folder, and the system picks them up and processes them unattended.
Yes. Romanian OCR is included, along with other languages, and mixed-language documents are handled in the same pipeline. Poor quality originals still get indexed, they simply have a higher chance of landing in the review queue.
That depends on your local retention rules and your accountant's advice. The archive gives you searchable PDFs with a full-text index, which covers day-to-day operations. Many companies keep paper for the legally required period and then reclaim the storage room.
There is a one time setup/implementation cost, which covers the software licence, the AI module, installation on your virtual machine, SMB and VPN configuration, document type definition and team training. Ongoing cost is 25 EUR per month, including software and security updates, three support tickets and 1.000 pages of AI document processing. Backfilling a large historical archive is quoted per page. Prices exclude VAT.
Yes. The stack is packaged as a Docker Compose deployment on an Ubuntu VM, so Proxmox, VMware, Hyper-V or a plain Linux server all work. A NAS that supports Docker can host it too, with the scan and archive folders mapped to shared folders on the device.
Thirty minutes is enough to review your document types, scanning volume and storage, and to agree on the pilot that makes sense. No slide deck required.