Formalbyte logoFormal
Byte
Secure Data Solutions
Self-hosted AI document management

Find any document in seconds. Keep every file on your own server.

Formalbyte e-Document Manager turns scanned paper into a full-text searchable electronic archive. Built on open source, classified by OCRskill agentic OCR, deployed on a virtual machine your company owns.

One time implementation fee, then 25 EUR per month. See what is included

  • Documents stay on your storage
  • Private AI hosted in the EU
  • GDPR agreement included
  • Romanian OCR included
  • OCRskill.com
archive.your-company.internal
acme invoice july 2026

3 results in 0.12 s

Invoice 4127 — Acme Distribution

Invoice14 Jul 20262026/Invoice/07-July/

Service estimate 8842 — Acme Distribution

Service estimate09 Jul 20262026/Service-estimate/07-July/

Delivery note 1190 — Acme Distribution

Delivery note03 Jul 20262026/Delivery-note/07-July/

1 document is waiting for human review. Nothing uncertain is filed on its own.

up to 95%

of documents filed automatically, without human input

seconds

to find any document by the words printed on it

zero

documents stored outside your own infrastructure

1 week

from kickoff call to a working archive, typically

Why now

The archive room is not storage. It is a liability.

Cardboard boxes labelled 2019, Invoices or, worse, Misc. One document lost during an audit, a warranty dispute or a GDPR request can cost more than the entire system below.

Paper archive today

  • Twenty minutes of walking, digging and dust to find one invoice
  • A document that cannot be found during a tax audit or a warranty dispute
  • Square metres of shelving paid for every single month
  • Only the colleague who built the filing system knows where anything is
  • Every scan labelled by hand: type, date, partner, folder

With e-Document Manager

  • Two words in the browser and the document opens in seconds
  • A full date range exported for the auditor while they wait
  • Shelves emptied and the archive room turned back into usable space
  • Every authorized colleague finds documents without asking anyone
  • AI fills in type, date and correspondent; people only check exceptions

Capabilities

Less filing. Faster answers. A cleaner audit trail.

Full-text search across every page

Search the content of every processed document, not just filenames. Image-only PDFs and photographed pages become searchable text through OCR.

Agentic AI classification

OCRskill reads each document and fills in document type, issue date and correspondent, up to 95% of the time without anyone typing a label.

Human review queue

Unreadable, unusual or out-of-scope documents are flagged for a person instead of being filed blindly into the archive.

A predictable archive structure

Files land in a consistent year, document type, month and correspondent path directly on your own storage.

Romanian and multilingual OCR

Diacritics and mixed-language documents are handled in the same pipeline, including scans made from imperfect originals.

Access control and audit trail

Individual accounts, permissions per document type and a complete history of who changed what and when.

Unattended intake from your scanner

Staff scan whole stacks into a shared SMB folder. The system collects the files and processes them without further clicks.

Works in any browser

Search and manage the archive from a normal web address over HTTPS. There is nothing to install on every laptop.

How it works

Scan the stack. Walk away. The archive organizes itself.

  1. 01 Scan a whole stack

    Staff feed documents into the office scanner, stacks at a time, and save them to the shared scans folder.

  2. 02 OCR reads every page

    The system picks up new files automatically and extracts the text, including from image-only PDFs.

  3. 03 AI classifies the document

    OCRskill determines the document type, the real issue date printed on the page and the correspondent.

  4. 04 Exceptions go to a human

    Anything unclear or unrelated to your archive is flagged for review rather than filed incorrectly.

  5. 05 The document becomes findable

    It is filed into the right folder and is instantly searchable for every authorized user in the browser.

Where the file lands

/archive/2026/Invoice/07-July/Acme-Distribution-Invoice-4127.pdf

Year, document type, month, correspondent. If you know what the document is, you already know where it lives. No more Misc folder, no more scanned_document_0002.pdf.

Data sovereignty

Your documents never leave your building

Confidentiality is the foundation of the architecture, not a setting you switch on later. The AI reads text through an encrypted tunnel and forgets it immediately.

  • The platform runs as a virtual machine on infrastructure you own
  • Documents are stored exclusively on your storage, shared over SMB
  • The AI module uses a private large language model hosted in the European Union
  • Document text travels through an encrypted VPN tunnel and is never stored on the AI server
  • No document ever reaches a cloud provider outside the EU
  • A complete GDPR storage and processing agreement is part of the delivery

Deployment stack

A maintainable Docker Compose deployment on an Ubuntu VM you provide, wired into the network folders your team already uses. No exotic dependencies, no vendor lock-in.

Ubuntu VMDocker ComposeOpen Source doc. managementOCRskill agentic OCRTraefik + HTTPSRedisApache TikaGotenbergSMB sharesPrivate EU LLMVPN tunnel

AI with a human checkpoint

Automation absorbs the repetitive metadata work. People keep the final say on uncertain or exceptional documents, so a wrong guess never becomes a wrong archive.

Two editions

Rules get you halfway. AI gets you to the finish line.

On an archive of 40.000 documents, the gap between 55% and up to 95% automatic classification is measured in thousands of hours of manual work.

Simple edition

Rules only, no AI module

Scanning, OCR and full-text search
Predictable folder structure
Automatic classification
around 55%, the rest by hand
Human effort per document
significant manual tagging
Best for
small volumes and a minimal budget
Recommended

AI edition

OCRskill agentic classification

Scanning, OCR and full-text search
Predictable folder structure
Automatic classification
up to 95%, depending on scan quality
Human effort per document
checking exceptions only
Best for
historical archives plus the daily document flow

Investment

One setup fee, one small monthly cost, no per-seat surprises

Initial Setup

One time cost

    Includes:

  • Software licence, including free security updates
  • AI classification module licence
  • Secure VPN tunnel to the EU AI datacentre
  • Installation of the Linux virtual machine on your infrastructure
  • SMB/Windows storage with folder auto-scan configuration
  • Access expert implementation consultant
  • Document types defined around how your business actually works
  • Training for the team that will operate the scanning
  • GDPR data storage and processing agreement

Maintenance and AI processing

25EUR / month

  • Software updates and security patches
  • Three support tickets per month with a 24h response target
  • 1.000 pages of AI document processing every month
  • Access to documentation and team training material

Indicative pricing, excluding VAT. Importing a large historical archive is quoted per processed page after a short scoping call, and you only pay for the pages actually processed.

What you get back

The cost of keeping the paper archive

20 min to 5 sec

One document lookup

A trip to the archive room becomes a search box. Multiply by every lookup your team makes in a month.

~1.900 hours

Manual filing avoided

Tagging 40.000 documents by hand costs roughly 2.000 hours. At up to 95% automatic classification, almost all of it disappears.

30+ boxes

Shelving reclaimed

80.000 scanned pages is more than thirty archive boxes you no longer need to store, dust or move.

days to minutes

Audit preparation

Instead of pulling binders for a week, export the requested date range and hand it over.

Illustrative figures for an archive of roughly 40.000 documents at about two pages each. Your own numbers depend on volume, scan quality and how often documents are retrieved.

A strong fit for

Built for teams that are drowning in paper

Accounting and finance teams

Invoices, receipts, purchase orders and bank documents, ready for any inspection.

Law firms and consultancies

Contracts, case files and client correspondence, searchable by any clause.

Car dealerships and service centres

Invoices, service estimates, warranty files and supplier correspondence.

HR departments

Employee files and administrative paperwork with permissions per document type.

Logistics and customs teams

Delivery notes, CMRs and transport paperwork found by shipment reference.

Audit and compliance heavy businesses

Clean, date-stamped archives and fast answers to GDPR or regulatory requests.

Rollout

Start with one category, then retire the filing cabinets

01 Scoping call

Thirty minutes to map document categories, users, storage and the current scanning flow.

02 Deployment

The stack goes live on your virtual machine, usually in under a week, storage and VPN included.

03 Pilot

Start with one category and one week of new documents, then confirm classification and folder paths.

04 Backfill and scale

Import the historical archive in controlled batches, train the team, retire the filing cabinets.

FAQ

Questions buyers ask before they sign

What is Formalbyte e-Document Manager?

It is a self-hosted electronic document archive built on open source and extended with OCRskill agentic OCR. It runs on an Ubuntu virtual machine on your own infrastructure, indexes every scanned page for full-text search, and fills in the document type, issue date and correspondent automatically so nobody has to type labels by hand.

Where are my documents actually stored?

On your own storage. The platform runs as a virtual machine on your infrastructure, and both the scan intake folder and the finished archive live on company storage shared over SMB. Formalbyte does not host your documents.

Is my data sent to a public cloud AI service?

No. The AI module uses a private large language model hosted in the European Union. Document text travels encrypted through a VPN tunnel, is classified, and the result returns to your system. Documents are never stored on the AI server and never reach a cloud provider outside the EU. A full GDPR storage and processing agreement is included.

How accurate is the automatic classification?

With the AI module we typically see up to 95% of documents classified correctly without human input, depending strongly on scan quality, language and how consistent your document types are. A rule-based setup without AI lands closer to 55%. Anything the system is unsure about is flagged for human review instead of being filed blindly.

How long does the deployment take?

Usually under a week after a 30-minute scoping call: installing the virtual machine, connecting SMB storage, configuring the VPN, defining your document types and training the team. The slow part is never the software, it is scanning the historical boxes.

Does it work with the scanner we already own?

Yes, as long as the scanner can save PDFs or images to a network folder. Staff feed whole stacks into a document scanner, the files land in the shared scans folder, and the system picks them up and processes them unattended.

Does it handle Romanian documents and diacritics?

Yes. Romanian OCR is included, along with other languages, and mixed-language documents are handled in the same pipeline. Poor quality originals still get indexed, they simply have a higher chance of landing in the review queue.

Do we still have to keep the original paper?

That depends on your local retention rules and your accountant's advice. The archive gives you searchable PDFs with a full-text index, which covers day-to-day operations. Many companies keep paper for the legally required period and then reclaim the storage room.

What does it cost?

There is a one time setup/implementation cost, which covers the software licence, the AI module, installation on your virtual machine, SMB and VPN configuration, document type definition and team training. Ongoing cost is 25 EUR per month, including software and security updates, three support tickets and 1.000 pages of AI document processing. Backfilling a large historical archive is quoted per page. Prices exclude VAT.

Can it run on our existing hypervisor or a NAS?

Yes. The stack is packaged as a Docker Compose deployment on an Ubuntu VM, so Proxmox, VMware, Hyper-V or a plain Linux server all work. A NAS that supports Docker can host it too, with the scan and archive folders mapped to shared folders on the device.

Bring one document workflow. We will map the path forward.

Thirty minutes is enough to review your document types, scanning volume and storage, and to agree on the pilot that makes sense. No slide deck required.