DEV Community

TankhaPay
TankhaPay

Posted on

Integrating an HRMS With an Existing ERP Without Replacing It: Sync Patterns That Actually Hold Up

Most companies that add an HRMS are not trying to get rid of their ERP. The ERP runs finance, procurement and inventory, and it holds years of audited history. Nobody wants to migrate the general ledger just because attendance tracking is weak.

The real gap is usually HR depth. ERP HR modules often cover the basics, an employee record and maybe a payroll run, but fall short on things like shift rostering, geofenced or biometric attendance, statutory payroll compliance across regions, expense workflows, performance cycles and headcount planning. So teams bring in a dedicated HRMS and keep the ERP as the financial system of record.

That leaves you with two systems that both care about employees, cost centres and money. This post covers the sync patterns that hold up in production, and the ones that quietly fall apart.

Step 1: Decide the source of truth for each entity

Integration problems are usually ownership problems. If two systems can both edit the same field, you will eventually get a conflict that no sync job can resolve. Before you write any code, agree on one owner per entity, and ideally per field.

A typical split looks like this:

Entity Usual owner Flows to Notes
Employee master (personal, job, status) HRMS ERP ERP only needs a subset: ID, name, entity, cost centre, status
Cost centres / departments ERP HRMS Finance controls the chart; HR consumes it
GL accounts ERP HRMS (as mapping targets) Never created from the HR side
Attendance, leave, shifts HRMS ERP (summarised, if at all) Raw punches rarely belong in the ERP
Payroll results HRMS ERP as journals Posted as summarised journal entries, not per-payslip
Vendor / bank master for reimbursements ERP HRMS (read-only) Avoid duplicate payee records

Write this table down and get finance and HR to sign off on it. Every later design decision depends on it.

Step 2: Pick the right transport per flow

There's no single best pattern. Each flow has its own volume, latency and failure tolerance.

Scheduled batch (files over SFTP, or bulk API exports). Good for payroll journals, monthly accruals and full-snapshot reconciliations. Batch is easy to audit: you have a file, a checksum and a timestamp. The downside is latency, plus the temptation to "just re-send the whole file", which leads to duplicates if the receiver isn't idempotent.

Synchronous REST APIs. Good for lookups and low-volume writes where the user is waiting, like validating a cost centre code when HR creates a position. Keep these calls narrow, and never let an HR screen block on an ERP that's down for month-end maintenance.

Webhooks / events. Good for near-real-time propagation of master data changes such as new joiners, transfers and exits. Events reduce polling, but they arrive out of order, arrive twice, or don't arrive at all. Treat every event as a hint to go fetch the current state, not as the state itself.

A pragmatic combination that works well:

  • Events for employee lifecycle changes (with a nightly full sync as a safety net)
  • REST lookups for reference data validation
  • Batch for anything that becomes a financial posting

Step 3: Make every write idempotent

If you remember one thing from this post, make it this: every write across the boundary must be safe to repeat. Retries, replayed webhooks and re-uploaded files are normal. Duplicated employees or double-posted payroll are not.

Use a stable business key (the employee ID owned by the source system), not the target system's internal ID. Carry an idempotency key on each change, derived from the source record and its version, and store the keys you have already processed.

Here's a minimal Python sketch of an idempotent employee upsert into an integration staging store that sits in front of the ERP:

import hashlib
import json
from datetime import datetime, timezone

SYNC_FIELDS = ["employee_id", "legal_entity", "full_name",
               "cost_centre", "status", "date_of_joining", "date_of_exit"]

def payload_hash(record: dict) -> str:
    subset = {k: record.get(k) for k in SYNC_FIELDS}
    return hashlib.sha256(json.dumps(subset, sort_keys=True, default=str).encode()).hexdigest()

def upsert_employee(conn, record: dict, source_version: int) -> str:
    idem_key = f"emp:{record['employee_id']}:v{source_version}"
    h = payload_hash(record)

    with conn:  # one transaction
        cur = conn.cursor()
        cur.execute("SELECT 1 FROM processed_keys WHERE idem_key = %s", (idem_key,))
        if cur.fetchone():
            return "duplicate_skipped"

        cur.execute("""
            INSERT INTO erp_employee_stage
                (employee_id, legal_entity, full_name, cost_centre, status,
                 date_of_joining, date_of_exit, source_version, payload_hash, updated_at)
            VALUES (%(employee_id)s, %(legal_entity)s, %(full_name)s, %(cost_centre)s,
                    %(status)s, %(date_of_joining)s, %(date_of_exit)s, %(v)s, %(h)s, %(ts)s)
            ON CONFLICT (employee_id, legal_entity) DO UPDATE SET
                full_name = EXCLUDED.full_name,
                cost_centre = EXCLUDED.cost_centre,
                status = EXCLUDED.status,
                date_of_exit = EXCLUDED.date_of_exit,
                source_version = EXCLUDED.source_version,
                payload_hash = EXCLUDED.payload_hash,
                updated_at = EXCLUDED.updated_at
            WHERE erp_employee_stage.source_version < EXCLUDED.source_version
              AND erp_employee_stage.payload_hash <> EXCLUDED.payload_hash
        """, {**record, "v": source_version, "h": h,
              "ts": datetime.now(timezone.utc)})

        cur.execute("INSERT INTO processed_keys (idem_key) VALUES (%s)", (idem_key,))
    return "applied"
Enter fullscreen mode Exit fullscreen mode

Two details matter here. The source_version guard stops an older, late-arriving event from overwriting newer data. The hash comparison skips no-op updates, which keeps downstream change logs quiet.

Step 4: Use change data capture instead of "modified since"

Polling WHERE updated_at > :last_run looks simple, but it misses records updated inside the same second, records whose clocks drift, and deletes. Where you can, use a proper change feed: an outbox table written in the same transaction as the business change, a database log-based CDC tool, or a vendor change API with a cursor or sequence number.

The outbox pattern works especially well on the HRMS side. Each employee change writes a row like (aggregate_id, version, event_type, payload, created_at), and a relay publishes those rows reliably. The version number then becomes your idempotency key on the other side.

Step 5: Mapping tables, not hardcoded codes

HR thinks in departments, locations and pay components. Finance thinks in cost centres, profit centres and GL accounts. Those lists don't line up one-to-one, and both change over time.

Keep explicit, versioned mapping tables in the integration layer:

  • cost_centre_map(hr_department, hr_location, legal_entity, erp_cost_centre, valid_from, valid_to)
  • gl_map(pay_component, legal_entity, employee_category, debit_gl, credit_gl, valid_from, valid_to)

Effective dating matters. When finance restructures cost centres mid-year, payroll for earlier periods must still post to the old codes. And when a lookup finds no mapping, fail loudly by routing the record to an error queue. Never fall back silently to a "default" cost centre.

Step 6: Multi-entity and multi-state realities

Groups with several legal entities, or operations across several states, add a dimension to almost every key:

  • Employee identity should be unique per legal entity, or globally with entity as an attribute. Decide which, and handle inter-company transfers as an exit plus a join, or as an explicit transfer event.
  • Statutory deductions and employer contributions can vary by state and by entity registration. Map each liability to the right entity's payable account, not a single group-level account.
  • Journals must balance per entity. A group-level balance can hide an entity-level imbalance.

Step 7: Posting payroll journals back to the ERP

Post summarised journals, not one entry per payslip. Aggregate by entity, period, cost centre and GL account, and include enough references to drill back into the HRMS.

{
  "idempotency_key": "payroll:ENT01:2026-09:run-2",
  "legal_entity": "ENT01",
  "period": "2026-09",
  "posting_date": "2026-09-30",
  "currency": "INR",
  "source_system": "HRMS",
  "source_run_id": "PR-2026-09-ENT01-002",
  "lines": [
    { "gl_account": "510100", "cost_centre": "CC-OPS-01", "debit": 1850000.00, "credit": 0, "memo": "Basic + allowances" },
    { "gl_account": "510200", "cost_centre": "CC-OPS-01", "debit": 222000.00, "credit": 0, "memo": "Employer statutory contributions" },
    { "gl_account": "220400", "cost_centre": null, "debit": 0, "credit": 395000.00, "memo": "Statutory deductions payable" },
    { "gl_account": "220100", "cost_centre": null, "debit": 0, "credit": 1677000.00, "memo": "Net salary payable" }
  ],
  "control_totals": { "line_count": 4, "total_debit": 2072000.00, "total_credit": 2072000.00, "headcount": 412 }
}
Enter fullscreen mode Exit fullscreen mode

The idempotency_key includes the run number, so a corrected re-run is a distinct posting while a retried upload of the same run is rejected. Control totals let the receiver validate before posting. If corrections are needed after posting, send a reversal followed by a fresh journal rather than editing posted entries.

Step 8: Reconciliation reports and error queues

Every integration drifts sooner or later. Plan to detect drift instead of hoping it won't happen.

Error queue. Any record that fails validation (missing mapping, unknown cost centre, unbalanced journal) goes to a durable queue with the payload, the error, the attempt count and an owner. Retry transient failures with backoff. Send business errors to a person, because retrying won't fix a missing GL mapping.

Reconciliation. Run scheduled comparisons between the two systems. Here is a simple headcount and cost reconciliation by entity and cost centre, assuming both sides are landed in a reporting schema:

WITH hr AS (
  SELECT legal_entity, cost_centre,
         COUNT(DISTINCT employee_id) AS headcount,
         SUM(gross_cost)             AS gross_cost
  FROM hr_payroll_results
  WHERE period = '2026-09'
  GROUP BY legal_entity, cost_centre
),
erp AS (
  SELECT legal_entity, cost_centre,
         SUM(debit - credit) AS posted_cost
  FROM erp_journal_lines
  WHERE period = '2026-09' AND source_system = 'HRMS'
    AND gl_account IN (SELECT debit_gl FROM gl_map)
  GROUP BY legal_entity, cost_centre
)
SELECT COALESCE(hr.legal_entity, erp.legal_entity) AS legal_entity,
       COALESCE(hr.cost_centre,  erp.cost_centre)  AS cost_centre,
       hr.headcount,
       hr.gross_cost,
       erp.posted_cost,
       COALESCE(hr.gross_cost, 0) - COALESCE(erp.posted_cost, 0) AS variance
FROM hr
FULL OUTER JOIN erp
  ON hr.legal_entity = erp.legal_entity AND hr.cost_centre = erp.cost_centre
WHERE ABS(COALESCE(hr.gross_cost, 0) - COALESCE(erp.posted_cost, 0)) > 1
ORDER BY ABS(COALESCE(hr.gross_cost, 0) - COALESCE(erp.posted_cost, 0)) DESC;
Enter fullscreen mode Exit fullscreen mode

The FULL OUTER JOIN is the important part. It shows cost centres that exist on only one side, which is where mapping bugs usually hide. Run a similar check for the employee master (active in HRMS but missing in ERP, and the reverse).

Step 9: Security and privacy

HR data is some of the most sensitive data a company holds. Treat the integration as its own security boundary:

  • Least privilege. Use a dedicated service account per flow. The journal poster can create draft journals in specific ledgers and nothing else. The master-data sync can't read salaries.
  • Minimise fields. The ERP rarely needs national ID numbers, bank details, addresses or salary components at employee level. Don't send what the target doesn't need.
  • Mask PII in logs and error queues. Payloads end up in logs, dead-letter queues and support tickets. Mask or tokenise identifiers before they get there.
  • Encrypt in transit and at rest. Use TLS for APIs, key-based SFTP with encrypted files, and short-lived credentials stored in a secrets manager, never in config files.
  • Audit everything. Keep a record of who or what changed each record, when, and the idempotency key behind it. Auditors will ask.

Where TankhaPay fits

At TankhaPay we built our enterprise-grade HRMS and AI recruitment platform to work for every workforce type, and to sit alongside the ERP you already run rather than replace it. Teams keep finance in their ERP while getting advanced HR capabilities: attendance with geofencing, face and biometric capture, shift and roster management, payroll with PF/ESI/PT/TDS/LWF/gratuity compliance, expense, asset and performance management, and headcount planning. If that's the gap you're filling, take a look at our HRMS that integrates with your existing ERP.

Wrapping up

Keeping your ERP and adding an HRMS is a sound architecture, provided you treat the boundary between them deliberately. Assign one owner per entity, choose the transport per flow, make every write idempotent, keep mappings explicit and effective-dated, post balanced and summarised journals, and reconcile on a schedule. Get those right and the integration fades into the background, which is where it belongs.

Top comments (0)