softmedia/ digital systems
PL
System online00:00:00
ROOT / CASE STUDIESSECURE CONNECTION
CASE_03 / DATA_QUALITY

DbFixStudio: deduplication and normalisation of a long-running customer database

DbFixStudio is a custom platform for cleaning, normalising and deduplicating business and customer data. It combines configurable data-quality rules, similarity methods, machine learning and operator control. The system detects likely duplicates, distinguishes branches of the same organisation, selects master records and prepares data for safe consolidation or further analysis.

Project type
ERP and CRM data / data quality / machine learning
Status
operational system under iterative development
Timeframe
MVP within weeks; approximately eight months for the current scope
Scope
DATA_QUALITY / DEDUPLICATION / ML.NET / SQL
01 / CASE_BLOCK

The challenge

Long-running customer and company databases rarely remain consistent. Records are entered by different teams, imported from several systems and maintained under changing rules. One organisation may appear under different names, addresses or identifiers.

Exact matching is not enough, while loose automatic matching can merge separate companies or branches. The labour-intensive part needed automation, with human control retained for ambiguous cases.

02 / CASE_BLOCK

A controlled data-quality workflow

  1. import data from SQL, files or another source
  2. map source columns to known data types
  3. normalise values with configurable rules
  4. generate credible candidate pairs
  5. calculate field-level similarity
  6. review pairs and record decisions
  7. train the model on approved decisions
  8. select master records using business rules
  9. export results for consolidation or analysis
03 / CASE_BLOCK

Normalisation without losing source data

Original values remain available for audit. Normalisation creates an additional comparison layer instead of irreversibly overwriting source records.

A dedicated address engine separates street types, names and building or unit numbers, handling inconsistent abbreviations and formats to reduce false matches.

04 / CASE_BLOCK

Machine-learning-assisted deduplication

Different fields carry different evidential value. A matching tax identifier is a strong signal, while a similar name without a matching address may represent another entity.

The operator marks pairs as duplicates, non-duplicates or uncertain. The model learns from approved decisions but never performs an irreversible merge by itself.

05 / CASE_BLOCK

Business outcome

DbFixStudio reduces the number of records requiring manual comparison and turns one-off database cleaning into a repeatable, auditable process.

Better normalisation reduces false matches, while human review limits the risk of combining separate entities. The resulting data can support migration, reporting and further analysis.

TX / STACK

Technologies

  • .NET and ASP.NET Core
  • SQL and SQL Server LocalDB
  • ML.NET
  • HTML, CSS and JavaScript
  • Address parsers and reference dictionaries
  • Configurable normalisation rules
  • Batch processing
  • AI/LLM-assisted pair assessment
Delivery and write-upSoftMedia
Author / responsible personPiotr Szymański
Last updated
READY_FOR_DISCOVERY

Let’s discuss a similar challenge.

Tell us about your project