The difficult part of collecting Avaya CDR was not parsing CSV or opening a TCP port. It was deciding when a record was durable, when a file was safe to remove, and when a call that had already been processed was no longer complete.
Communication Manager can send another record for the same UCID after the call is already visible. Equinox can change an XML file without changing its name. SBCE can export several rows for one session, with useful fields distributed unevenly between them. Unformatted CM records may have no reliable correlation identifier at all.
Avaya documents and exports CDR separately for each component. Session Manager stores its own records; SBCE sends records to an external CDR adjunct by SFTP or RADIUS; CM and Equinox expose different formats again. We found no vendor-supplied tool that ingests CM, SM, SBCE, and Equinox into one reprocessable correlation model. That was the gap we needed to fill.
The implementation and all scenarios described below were tested with Avaya Aura 8.1 and Equinox 9.1.12.
The requirements
The collection path had to satisfy five rules:
- Do not acknowledge data before it is stored durably.
- Do not remove a source file before its RAW transaction commits.
- Make retries idempotent.
- Rebuild an existing call when a late correlated record arrives.
- Preserve source records so processing logic can be changed and rerun.
That led to this data path:
CM TCP / SM SFTP / EQ rsync / SBCE SFTP
|
v
source-specific collectors
|
v
Unix socket typed-batch protocol
|
v
SQLite WAL durable spool
|
v
PostgreSQL source RAW tables
|
v
source-specific processors
|
v
Calls + Details + UI/API views
The SQLite spool is not a cache. It is the first durability boundary. It runs in WAL mode with synchronous=FULL. A collector receives acceptance only after the local spool transaction commits. The spool removes a batch only after PostgreSQL commits the corresponding RAW transaction.
This keeps a PostgreSQL outage, process restart, or network interruption from turning an accepted batch into missing data.
Why we did not build one universal processor
The four sources describe different things:
- CM provides call legs and condition codes;
- SM provides session records;
- SBCE provides border-session records;
- Equinox provides conference and event XML.
Forcing them into one correlation model would create plausible but false call chains. We use a common transport and storage pipeline, then source-specific processors and Details tables.
The UI can present a common operational surface without pretending that a CM UCID, an SBCE Session Id, and an Equinox event have identical semantics.
The CM issue: a processed call can change
Our initial assumption was simple: process pending RAW rows, create the call, and mark the rows complete.
That assumption fails for transfers, conferences, and forwarding. Another CM record with the same UCID can arrive after the existing call has already been processed. Appending only the new row loses the original context. Creating another call duplicates the session.
The current rule is:
- group by UCID only when UCID is present and sequence number is numeric;
- when any pending row is found, reload every RAW row for that UCID, including processed rows;
- rebuild the ordered chain;
- UPSERT the Calls row by
call_key; - replace its derived details in the same transaction.
If UCID or a valid sequence number is missing, the row remains independent. We deliberately do not correlate by telephone number and time window. Twenty simultaneous calls from one number are valid traffic, not evidence of one call chain.
The main row stores the first relevant source party and final relevant destination. Intermediate transitions and the ordered condition-code sequence remain in Details.
The SBCE issue: several rows describe one session
SBCE correlation is scoped by source server. The key priority is:
-
Session Id; - UCID;
- internal call key;
-
Call Id; - RAW row identifier.
Grouping the rows is only half of the problem. Selecting MAX(duration), MAX(routing_profile), and MAX(disconnect_reason) independently can construct a call that never existed.
Instead, the processor selects one representative row using a deterministic score:
- duration;
- presence of Disconnect Time;
- presence of UCID;
- source record number;
- RAW identifier.
Parties, Routing Profile, Server Flow, and disconnect reason then come from that same row. Other rows stay available in Details.
File collection is part of the transaction model
All remote transfers use a temporary .part file. The final local name appears only after a complete transfer.
Session Manager
CDR_User retrieves the remote CDR file through SFTP into a .part file. Only the complete file is parsed and passed to SM RAW.
Equinox Management
Equinox compares source server, filename, record number, and SHA-256. Changed XML is parsed again; an unchanged hash stops duplicate processing.
Source cleanup is conservative:
- SM and Equinox remote files are never modified or deleted;
- SBCE incoming files are removed locally only after RAW commit;
- CM Survivable collection ignores active
C-files and reads closed archives; - CM remote deletion targets only the exact stable file after RAW commit and copied-file registration.
Transfer success is not processing success. Operations must check the pending RAW queue as well as collector logs.
The second durability boundary
A processor writes the derived call, its source-specific details, and RAW completion state in one PostgreSQL transaction:
BEGIN;
-- UPSERT call
-- replace derived details
-- mark contributing RAW rows stat = true
COMMIT;
If any step fails, stat remains false and the batch is available for another attempt. There is no state where RAW is marked complete but the corresponding Calls row was not committed.
One operational trap: SBCE SFTP paths
The host path and chroot path are not interchangeable:
Host filesystem: /opt/cdr/data/sbce/1
SFTP/chroot: /sbce/1
SBCE Location: /sbce/1
CDR UI Path: 1
With an empty SBCE Location, the sender may try to open //S000...csv in the chroot root and receive Permission denied. Successful authentication proves the account works; it does not prove that the configured upload directory is writable.
Runtime and failure handling
The installation uses two Docker containers:
-
cdr-app: API, collectors, processors, maintenance, and UI; -
cdr-db: PostgreSQL.
Application processes are supervised inside cdr-app. Diagnostics report process state, load-balancer socket state, PostgreSQL health, pending RAW queues, ACD tables, disk usage, and recent supervisor events.
Recovery follows stored state instead of trying to infer what probably happened:
- collectors reconnect and retry;
- committed SQLite batches survive application restart;
- PostgreSQL RAW rows remain pending until processing commits;
- late correlated rows rebuild the derived call;
- original RAW data remains available for recalculation after processor changes.
What we deliberately do not solve
- Missing UCID cannot be reconstructed reliably from numbers and timestamps.
- Unformatted CM data without a stable identifier remains independent.
- SIP headers absent from SBCE CSV cannot be recovered later.
- CDR identifies signaling events; it is not a packet capture.
- Database insert benchmarks are not end-to-end call-capacity measurements.
These limits are preferable to confident but incorrect correlation.
Distribution
The public repository contains Docker release assets, installation and administration documentation, screenshots, and test reports. The application source is maintained privately.
- Repository: https://github.com/vovan-T/CDR-AVAYA
- Installer: https://github.com/vovan-T/CDR-AVAYA/releases/latest/download/install.sh
AI assistance was used for editorial structure and language review. Architecture and behavior were verified against the implementation and project documentation.






Top comments (0)