Data Cleaning and Dataset Preparation for AI
We turn accumulated CRM, 1C, Excel, and documents into a working foundation for AI: audit, cleaning, enrichment, and assembly of an AI-ready dataset or knowledge base for RAG and assistants — in 2–6 weeks.
- МСБ
- Крупный бизнес
- Госсектор
- Стартапы
Без обязательств: сначала разбираем ситуацию, затем предлагаем формат работы.
What You Get
- AI data audit in 3–7 days. We inventory all sources, calculate a Data Quality Score, identify risks under 152-FZ, and honestly tell you which scenario will pay off first — cleaning, RAG, analytics, or fine-tuning. This is the first step to avoid overpaying.
- Cleaning CRM, 1C, and spreadsheets. We remove duplicates, normalize names, phone numbers, addresses, INN, and dates. We restore the machine-readable form in which AI stops making errors at the input. Disputed merges go to you for confirmation — not "as we decided."
- Knowledge base for RAG from your documents. OCR for scans, text extraction from PDF/DOCX, removal of outdated versions and junk. The team stops searching through archives — the assistant answers according to regulations and FAQs.
- Data enrichment. We add attributes from external sources: OKVED, company status, geo, addresses. This strengthens both analytics and search accuracy.
- Structure for the target AI scenario. We deliver a Markdown corpus, JSONL dataset, CSV/SQL, or RAG-ready structure with metadata, 300–800 token chunks, and a golden set. Not an "archive of files," but a working loop.
- Quality control. Golden set, retrieval accuracy checks, and metrics of "completeness > 95%, chunk duplicates < 2%." You get a report, not promises — what it was and what it became.
Results in Numbers
These are verifiable benchmarks for typical loops (data volume and format affect the final estimate — we fix it before starting):
| Stage | Price "from" | Timeline | Deliverable |
|---|---|---|---|
| AI data audit | 50–150 thousand ₽ | 3–7 days | source map, Data Quality Score, legal check, roadmap |
| CRM / 1C / spreadsheet cleaning | 100–500 thousand ₽ | 1–4 weeks | deduplication, normalization, "before → after" report |
| Knowledge base for RAG | 250–900 thousand ₽ | 2–6 weeks | OCR, Markdown corpus, chunks, metadata, golden set |
| RAG assistant | 400 thousand – 1.5 million ₽ | 4–10 weeks | working loop for your data |
| LoRA / fine-tuning | 700 thousand – 3 million ₽ | 6–14 weeks | fine-tuning for style, classification, or industry |
Benchmark for a small pilot: cleaning a CRM of 10,000 records comes to approximately 200,000 ₽ with transparent decomposition by people, API, and timelines.
Entry rule: we almost always start with an audit — it's inexpensive and reduces the risk of unnecessary spending. A sample of 100–1,000 rows or 20–50 documents is enough to assess the scope.
Why Us
- We start with an audit, not with selling a model. In 3–7 days we'll show you what's ready for AI, where the 152-FZ risk is, and which scenario will pay off first.
- We work within the Russian legal framework. We account for 152-FZ, commercial secrecy regime, depersonalization, and restrictions on transferring data to external cloud APIs. Data doesn't "leak abroad" without your consent.
- Not a "black box." The output includes a README, quality report, metrics, versioning rules, and acceptance criteria. You know what was done and how to reproduce it.
- We work with both spreadsheets and documents. In a single project we clean up CRM, 1C, Excel, PDF, DOCX, and file storage.
- Markdown-first where it improves quality. For documents we build a Markdown corpus: it preserves structure better, is easier to chunk, and simpler to version.
- We don't push fine-tuning where it's not needed. For most SME clients, it's more sensible to start with RAG and a clean knowledge base, and only add fine-tuning when there's a real need.
- Vendor-independent advice. We recommend the solution for your task, not the one that's more profitable for us.
Who It's For
For whom this quickly solves the problem:
- SMBs without a data team — have CRM, 1C, Excel, and manual processes; need a first AI scenario without hiring engineers.
- Startups — have an AI product idea but no clean dataset; need to quickly assemble a working set for an MVP or RAG.
- Large companies — pilots already exist, but data across departments is fragmented and not ready for scaling.
- Government agencies and regulated industries — personal data, regulations, and localization within the Russian framework without violating requirements.
When This Isn't Your Case
- Data is completely unstructured and unsystematized. If it's a "box of papers" without a system, you first need basic digitization, not AI preparation — we'll honestly discuss this at the audit.
- You're not ready to share even a sample. A minimum of 100–1,000 rows or 20–50 documents is needed to assess scope and risks. Without it, we won't guess.
- You don't actually need AI yet. Then we'll say it directly: don't waste your budget. Sometimes the right answer is not to launch the project.
Where to Start
The main step is a free AI data audit. In a short consultation we clarify the scope, sources, and personal data risk before you pay. You get decision-making tools: what's actually ready for AI, which path will pay off faster.
The secondary step is a demo on your sample: we'll show your Data Quality Score on 100–1,000 rows and what the result looks like, without verbal promises.
Then down the chain, if the data is ready, we move to AI consulting — so you don't stop at a clean base, but reach a working assistant, RAG, or analytics through a plan with ROI.
[Request a free AI data audit] · [See how we calculate the Data Quality Score] · [AI consulting — the next step]
FAQ
Where's the best place to start if we're just thinking about AI?
With an AI data audit. It shows which sources already exist, where the main quality problems are, and which scenario will deliver the fastest results.
Do you work only with spreadsheets or do you handle documents too?
Both types. We clean and normalize CRM, 1C, Excel, and SQL exports, and we build a corpus from PDF/DOCX/scans: OCR, text extraction, junk removal, working structure.
What's better in most cases: RAG or fine-tuning?
For most SME clients, it's more sensible to start with RAG: if the task involves documents, regulations, and FAQs — it launches faster, is cheaper, and is easier to update. We add fine-tuning for style, classification, or deep adaptation.
What if the data contains personal data or commercial secrets?
We account for this at the audit stage: NDA, personal data processing agreement, depersonalization, work in the client's secure environment or an agreed Russian infrastructure.
How much data is needed for the project to make sense?
A sample of 100–1,000 rows or 20–50 documents is enough for assessment. For RAG — from 100+ meaningful documents, for fine-tuning — an instruction dataset of 200+ quality pairs.
What most affects the cost?
The volume and format of data, the share of duplicates and errors, the need for OCR, the volume of manual validation, the presence of personal data, and whether you need just a clean dataset or a full knowledge base for RAG/fine-tuning.
Do you deliver only the dataset or take it to a working solution?
We can cover the full cycle: after preparation, we move the project into RAG, an AI assistant, LoRA/fine-tuning, or ongoing support mode.
---
Не нашли ответ? Обсудить задачу
Regulatory Context
- 152-FZ "On Personal Data" — lawful processing framework, access control, localization in the Russian Federation, depersonalization when necessary.
- Commercial secrecy — NDA, restriction of access, secure storage, prohibition on transferring to external services without consent.
- Cloud APIs and external models — we check consents, the provider's policy, and whether on-prem or a Russian cloud is needed.
Обсудим вашу задачу без лишнего риска
На первом шаге разберём ситуацию, покажем, где эффект, а где риски, и предложим формат работы под ваш масштаб.
Ответим в течение 1 рабочего дня. info@right-digits.ru