Reading as Everyone
FAQ

Straight answers, including the ones that are no.

Grouped by who usually asks. Pick a role and your group moves to the top. Every answer here is consistent with what the product does today — roadmap items are labelled as such.

General

What it is
01What is KognitCapture, in one paragraph?

Software you install on your own servers that turns incoming documents — scans, PDFs, images, office files, e-invoices — into checked, structured data. It imports from fourteen kinds of source, reads pages with OCR and optionally AI models, classifies and splits documents, extracts fields through versioned templates, validates them, sends the doubtful ones to a person, and exports the result to your systems.

02How is it different from a cloud capture API?

Three ways. Your documents are processed on your infrastructure, not sent to a provider. You pay a fixed annual licence instead of a price per page. And you get the whole process — review screens, workflow, validation, export — not only a recognition call.

03How is it different from a classic capture suite?

No page-based licence and no licence server. A browser interface for everything, including scanning. An API for everything the interface does. Language models as an option at several stages. And configuration by blueprints and a visual designer rather than a long consulting project per document type.

04How accurate is it?

We publish no accuracy percentage, because one number across unknown documents means nothing. Clean printed text reads well with Tesseract; poor scans, unusual typefaces and handwriting need a vision model, a trained language pack or a reviewer. The honest way to answer is to run your documents — which is what we offer.

→Can we use AI without sending anything to a cloud provider?

Yes. Text models in GGUF format run inside the application, and vision models run in a local model server managed by the platform on the same host. You can also point a step at your own Ollama, vLLM, LM Studio or other OpenAI-compatible server. In all of these cases pages stay on your network and there is no per-token charge — you pay in hardware instead.

Remote providers remain available per organisation and per step when you decide a hosted model is worth it, with a monthly spending ceiling.

→Is open-source OCR really good enough?

On clean 300 DPI print, Tesseract 5 is competitive. On degraded scans, dense forms and non-Latin scripts, the best commercial engines still read measurably better — we will not pretend otherwise. The answer for those documents is the hybrid path: Tesseract first, and low-confidence words or whole pages re-read by a vision model, local or remote, at your choice.

→What does a bake-off involve?

A fifteen-minute scanner and volume call first. Then someone reads your documents, authors templates, runs the Tesseract, hybrid and model paths over them, and writes up field-level accuracy with the cost per thousand pages of each mode. It runs one to two days — two is the ceiling. You keep the report whatever it concludes.

→Can I see it before talking to anyone?

Yes. The live demo is the application's own interface running in your browser on synthetic documents — about forty screens, no sign-up.

→What if the vendor disappears?

KognitCapture is self-hosted: there is no licence server that can switch your installation off. Source access and escrow are agreed in the contract conversation. Your documents are files on your disks and your data is in your own PostgreSQL database throughout.

05Which languages does the interface come in?

The interface is in English. Recognition is a different matter: around 129 OCR language packs can be installed, several per document.

06Who makes it?

KognitCapture is a product of Automatize BV in Heist-op-den-Berg, Belgium, developed together with A4IT and A4ME. It is designed, built and maintained in Europe.

Finance & accounts payable

Invoices
01Can it read invoice lines, not just header fields?

Yes. Table extraction uses column zones, row grouping and header detection, and follows a table across pages. Line items are also extracted by a model into the template's own row structure when you enable one. Complex line layouts are where a test on your suppliers' invoices tells you most.

02We have hundreds of suppliers with different layouts. Is that a template each?

One template with alternatives. Each alternative is a layout variant; all of them produce the same field set. For the long tail of suppliers you see twice a year, label search and model-based extraction read invoices without a dedicated variant.

03Does it handle Peppol and other e-invoices?

On the receiving side. It reads UBL and CII documents with an embedded PDF, and ZUGFeRD, Factur-X and XRechnung files, and maps them onto the same template as your paper invoices. It does not send e-invoices and is not a Peppol access point.

04Can it check that an invoice is correct?

It can check what rules can express: totals add up, VAT is consistent, the supplier exists in your master data, the purchase order number is found in a dataset, database or API. Whether the goods were received is a question for your ERP.

05Does it book into our accounting package?

It delivers the data in the format and through the channel your package accepts: a file, a database insert, a REST call. There are no named connectors for specific accounting products.

06Can an auditor see where a value came from?

Yes. Each field keeps the recognised value, every amendment, who verified it and when, next to the page image it was read from.

Operations & capture team

Daily work
01Do we have to install anything on the scanning PCs?

No. The Scan Hub runs in the browser and the server talks to network scanners that support eSCL (AirScan / Mopria).

02Our scanners use TWAIN or ISIS drivers. Will they work?

Not directly; those drivers are not supported by design. Set the scanner or its software to scan to a folder and let a hot-folder source import the files.

03What happens if the browser closes in the middle of a batch?

Nothing is lost. The scan session is held on the server; reopen it and continue.

04Can two operators end up in the same document?

No. Opening a document checks it out. Locks can cover a whole batch, expire by themselves, and can be released by an administrator.

05Can we assign documents to a specific reviewer?

Not to a named person today. Access is by hub grant per project, and the queue is ordered by service level, priority or arrival. Assignment to individual reviewers is not available.

06How do we know if we are meeting our deadlines?

Set service-level targets per project. The hubs can order work by deadline, documents close to breach are marked at risk, and alerts and digests are mailed.

07Can it read handwriting?

Only through a vision model; there is no separate handwriting engine. Expect handwritten fields to need review.

IT & infrastructure

Running it
01What do we need to run it?

A Linux x86-64 server, Java 25, and PostgreSQL. As a floor: 4 GB of heap for a small installation, 8 GB for real throughput. Local language models need additional memory outside the heap — budget roughly 4 to 5 GB for each 7-billion-parameter model at 4-bit quantisation.

02Windows or macOS?

Linux x86-64 is the supported platform, because that is where the bundled OCR libraries are self-contained.

03Containers? Kubernetes?

We do not publish images or manifests today. The deployment unit is a JAR run as a system service.

04How does it scale?

Add nodes. They share one PostgreSQL database and one filesystem; there is no message broker. Each node is given duties and capabilities, so you can dedicate machines to OCR or to export.

05Does it need a GPU?

Not for OCR, classification or the workflow engine. Local language models run on CPU and benefit from a GPU; whether you need one depends on the model and the volume.

06How do upgrades work?

Replace the JAR and restart. Database migrations run on start-up, and the built-in self-verification scenarios check the installation afterwards.

07Is there a native binary?

A GraalVM native-image build profile exists but is not validated, so we do not offer it as a supported deployment.

08What about backups?

Back up the PostgreSQL database and the file tree with your own tooling. The product has no built-in backup.

Developers & integrators

Building on it
01Is the whole product available through the API?

The web interface is itself a client of the REST API. The OpenAPI 3 specification is generated from the running application, and the built-in console shows the operations your account may call.

02How do integrations authenticate?

With a bearer token for the general API. API keys are used for specific lanes: embedded hubs, mobile devices, dataset loading and the MCP server.

03Will a template change break my integration?

Not if you pin a version. Publishing takes a snapshot; a workflow step or a capture call can name the version it wants.

04Are there webhooks or event streams?

Outbound: a webhook export goal with basic, bearer or API-key authentication, and batch completion callbacks. Inbound: a webhook import source verified by HMAC. There is no WebSocket or server-sent event stream.

05Can workflows run steps in parallel?

No. A document follows one path at a time, with decisions. Parallel branches and joins are not supported.

06Which model providers can I plug in?

Fifteen named providers, any OpenAI-compatible endpoint, local GGUF models, and anything else through an LLM-provider plugin — there is a sample project for exactly that.

Security & compliance

Sign-off
01Is KognitCapture GDPR compliant?

Software is not compliant; an organisation's processing is. What KognitCapture gives you is control: processing on your own infrastructure, access control, audit logs, retention policies and deletion. The documentation maps these controls to GDPR obligations to support your own assessment.

02Do you hold ISO 27001 or SOC 2?

No. There is no vendor attestation for KognitCapture. Because you host it, the software sits inside your own certified scope.

03Is there multi-factor authentication or single sign-on?

Not yet. Both are on the roadmap. Today users sign in with a username and password, protected by lockout and rate limiting.

04Are documents encrypted at rest?

Secrets — credentials and provider keys — are encrypted by the application with AES-256-GCM. Document files and the database are not; encrypt the volumes and the database with your platform's facilities.

05What is sent to an AI provider?

Nothing by default. If you configure a remote provider and use it in a step, the pages scoped for that step are sent, without automatic redaction. A local model keeps everything inside the installation.

06How long are documents kept?

Until you say otherwise. Retention policies are defined per project and are off until you create one.

More on the security page →

Buying

Management & procurement
01What does it cost?

One annual licence per installation, based on your turnover: from €900 a year, €2,500 at €5M turnover, €13,000 at €100M. Unlimited users and pages. The pricing page has the full table and a calculator.

02What else will we spend?

A server and database to run it on, the effort to configure templates and integrations, and — only if you choose a remote AI provider — that provider's usage charges.

03How locked in are we?

Your documents are files on your disks and your data is in your PostgreSQL database, in a documented schema. There is no runtime licence gate. Configuration can be exported, and a full project report documents every setting.

04How fast can we be live?

Installation is one JAR and a database. A blueprint gives you a running project in one step. The time that matters is tuning templates and rules on your documents and connecting the export — which depends on the documents.

05Can we see reference customers?

We prefer to prove it on your own documents. Ask us about references in the conversation.

06How does it compare with ABBYY, Kofax, Rossum or Textract?

There is a page for each of nine platforms, with both columns: where that product beats us and where we win, plus reported cost. Start at Alternatives, or put your page volume into the cost comparison.

07When is a metered service cheaper than this?

At low volume. Below roughly 100,000 to 150,000 pages a year a pay-per-page service is often cheaper than running your own infrastructure, and we will say so. Above that, self-hosting pulls ahead and the gap widens with every additional page.

Partners, ISVs & capture bureaus

Building a service on it
01Can we serve many clients from one installation?

Technically yes: tenants and organisations separate clients, with their own storage, branding, domains and mail providers. Commercially this is agreed separately from the standard licence.

→Is there a partner programme?

Yes, for resellers, software vendors embedding capture, and service providers running many clients. See Partners.

02Can we put our own brand on it?

Yes. Product name, logo, colours, typeface, sign-in page and footer links can be set at platform, tenant and organisation level.

03Can we embed it in our own software?

The Verification Hub and Scan Hub embed in an iframe, opened by a server-side session. Everything else is available through the API.

04Can we add our own connectors?

Yes, as plugins: import connectors, export connectors, workflow steps, email providers and model providers.

Question not here?

Ask it. An engineer answers within one business day.Ask us →
One question before you read

What is your role?

An accounts payable lead and a platform engineer need different answers. Pick a role and every page puts what matters to you first. Nothing is hidden for good, and you can change it at any time.

08Become a partner Sell and implement KognitCapture for your own clients. Opens the partner programme.
Stored only in this browser. No account, no tracking, no cookie.
Machine translation

Read this site in your language.

The site is written in English. Google Translate can show it in the languages below. The translation is automatic and not reviewed by us: for prices, licence terms and legal text the English page is the one that counts.

Translation by Google. Choosing a language loads Google Translate, sends the text of the pages you open to Google and stores one cookie in this browser. English switches all of that off. What is sent and stored