SourcepagePavel Novikau · private AI search over archives · self-hosted LLMs
[email protected]

Independent engineer · EU timezone · EN / RU / BE / PL

Your documents know more than your search box does.

I build private AI search and assistants over messy document collections — scanned, handwritten-era, multilingual, sensitive — running on your hardware or a server you control. Nothing leaves your perimeter unless you decide it should.

›wedding songs from Polesie villages
BEnewspaper scan · OCRp. 3
…запісаныя ў вёсцы вясельныя песні паўднёвага Палесся…
PLjournal article · born-digital
…pieśni weselne z pogranicza poleskiego…
RUfield notes · re-OCR
…свадебные песни сёл Полесья…

Illustration: one English query, matches in three languages — no translation step, no keyword lists. This is how my archive search engine works.

Who this is for

Teams sitting on documents nobody can search properly.

Archives, libraries, museums

Digitised collections that are online but effectively invisible — bad OCR, mixed languages, no good finding aid.

Law, consulting, due-diligence

Sensitive files that can't go to a public AI API, but need question-answering with citations.

Researchers & publishers

Literature, press and source corpora where you need the page, not a confident summary.

Small companies

A local LLM on your own box for internal docs and drafting — predictable cost, no data leaving.

Offers · fixed price

Start small, pay for a result, decide after.

Every engagement starts with the audit. If the prototype doesn't earn its keep on your own documents, you stop there and keep the report.

Step 1

Document AI audit

fixed €1,500

  • Sample of up to 1,000 of your documents
  • OCR & language quality assessment
  • Working search prototype on the sample
  • Written plan: stack, hardware, cost, risks

1 week · credited against step 2

Step 2 · most requested

Private archive search

from €6,000

  • OCR / re-OCR pipeline for scans
  • Multilingual semantic search with reranking
  • Answers that cite the exact source page
  • Access control, rate limits, cost caps
  • Deployed on your server or a private VPS

3–5 weeks · includes handover docs

Alternative step 2

Local LLM box

from €3,500

  • Model selection tested on your tasks
  • GPU server setup (consumer or workstation cards)
  • Chat + document Q&A for your team
  • Speech-to-text, translation on request

2–3 weeks · your hardware or mine to spec

After launch: support & improvements on a monthly retainer, or ad-hoc at €90/h.

Built & operated

Not slideware. Systems I run in production, on my own infrastructure.

SEMANTIKORAarchive search · RAG

Semantic search over Eastern European archives

Historical newspapers, journals and books — scanned and born-digital — searchable by meaning across languages. Research mode writes answers grounded in the retrieved sources. Invite-only perimeter with SSO, token validation, rate limits and a spend cap.

~45kdocuments
5+languages, cross-lingual
self-hostedembeddings, reranker, vector DB
PALEO-EUROPE 3Dgeodata · 3D · research

26,000 years of Europe's landscape, in the browser

Terrain, sea level, ice, vegetation, archaeological sites and ancient-DNA ancestry fields from the last glacial maximum to today — fused from public scientific datasets into an interactive 3D timeline.

175,733terrain tiles at z9
48climate epochs
49ancestry time slices
AGENT DEPARTMENTAI agents · automation

A ticket goes in, a reviewed pull request comes out

An autonomous pipeline: a task in the tracker is picked up by coding agents, cross-checked by a critic, and delivered as a pull request to a self-hosted Git server for a human to merge. Plus a Telegram project-manager agent with daily digests.

ClickUp→Gitend to end
dual + criticquality gate
humanalways merges
VAKOLICA.ORGmaps · generated content

Local history for every settlement in Belarus

An interactive map where each place gets an article built from retrieved sources. Includes a locally fine-tuned writer model evaluated for one thing above all: not inventing history.

vakolica.org ↗
1,100+pages in sitemap
0.97date grounding score
localLoRA writer model

All of it runs on infrastructure I built and maintain: a Proxmox cluster with GPU passthrough, dozens of containers, zero-trust remote access, monitoring and backups.

How a project runs

Short loops on your real documents.

  1. Audit

    You share a sample. I measure OCR quality, languages and what people actually search for.

  2. Prototype

    A working search on the sample. You try your own hardest queries on it.

  3. Build

    Full ingestion, tuning of retrieval on your query set, access control.

  4. Handover

    Deployment, runbook, a smoke test that proves search returns real sources.

Questions

The things people ask first.

Does our data go to OpenAI or anyone else?

Not by default. Embeddings, search and a local LLM all run on your side. A cloud model is an option you switch on per task, never a hidden dependency.

Our scans are bad. Is that a problem?

It's the normal case. Re-OCR, layout handling and noise-tolerant retrieval are a core part of the job, not an afterthought.

What hardware do we need?

Often less than expected: my own production stack runs on two consumer GPUs. The audit gives you a concrete spec and cost.

How do you stop the AI from making things up?

Answers must cite the source page, and I measure grounding on your data before launch — a fluent answer without a source counts as a failure.

Tell me about your documents.

Three lines are enough: what the collection is, roughly how big, and what you wish you could ask it. I reply within two working days.

First call is 30 minutes and free. The audit is where paid work starts.