This guide builds that summary in Python. It reuses the format-checking reader from Part 1, keeps rejected lines visible and sorts accepted groups by count. The result is a starting point for investigation, not a list of attackers or proof of complete network activity.
Start with the same documented input
The practice format has four fields separated by a tab: time, source address, destination address and destination port. Each line represents one invented connection-log event, not a packet or a complete session. It is a teaching format created for this guide, not a claimed Wireshark, Zeek or firewall export format.
The first line must be Time, Source, Destination, Port, with actual tabs between the names. The time uses exactly YYYY-MM-DDThh:mm:ssZ; Z means Coordinated Universal Time (UTC), the shared reference time used here rather than a local clock. Fractional seconds, time differences such as +01:00 and leap-second values are not supported. The addresses use Internet Protocol version 4 (IPv4), the four-number address format such as 192.0.2.10. Each part is separated by a dot, with no extra leading zeros. Internet Protocol version 6 (IPv6), the newer address format that uses groups such as 2001:db8::1, is outside this example. Ports are whole numbers from 1 to 65535; that is this format's limit, not a claim that port zero never appears elsewhere.
Use only the invented sample. Its addresses come from the ranges reserved for documentation by RFC 5737, the published internet reference. No real traffic was captured or analysed.
Create practice.tsv with the following actual tab-separated text:
Time Source Destination Port
2026-09-30T09:00:00Z 192.0.2.10 198.51.100.20 443
2026-09-30T09:00:01Z 192.0.2.10 198.51.100.20 443
2026-09-30T09:00:02Z 192.0.2.11 203.0.113.30 53
2026-02-30T09:00:03Z 192.0.2.10 198.51.100.20 443
2026-09-30T09:00:04Z 192.0.2.11 203.0.113.30 70000
Use the complete LogReader.py module included below in the same folder. It accepts this one teaching format, preserves physical line numbers and original text in memory, and rejects unsupported data lines with a reason. It stops the overall load for a missing file, unsupported header or exceeded file/row limit.
The article includes that exact module so the example can be reproduced without a different library or parser. It is not a separately validated reader for real Zeek, Wireshark or firewall exports. Do not feed a real export into it and interpret rejection as a problem with the network.
Save the shared reader
Save this as LogReader.py. It is the same module used in Part 1; both articles include it so this guide can be run on its own.
from dataclasses import dataclass as DataClass
from datetime import datetime as DateTime
from pathlib import Path
@DataClass
class LogEntry:
Number: int
Original: str
EventTime: str = ""
Source: str = ""
Destination: str = ""
Port: int = 0
Problem: str = ""
Accepted: bool = False
def Digits(Value, Maximum):
if not Value:
raise ValueError("empty number")
if any(Character not in "0123456789" for Character in Value):
raise ValueError("number needs ordinary digits")
Number = int(Value)
if Number > Maximum:
raise ValueError("number outside supported range")
return Number
def CheckAddress(Value):
Parts = Value.split(".")
if len(Parts) != 4:
raise ValueError("expected four IPv4 address parts")
for Part in Parts:
if len(Part) > 3 or (len(Part) > 1 and Part.startswith("0")):
raise ValueError("use canonical dotted IPv4 addresses")
Digits(Part, 255)
def CheckTime(Value):
if len(Value) != 20:
raise ValueError("expected time YYYY-MM-DDThh:mm:ssZ")
if any(Value[Index] != Mark for Index, Mark in
((4, "-"), (7, "-"), (10, "T"), (13, ":"), (16, ":"), (19, "Z"))):
raise ValueError("expected time YYYY-MM-DDThh:mm:ssZ")
NumberText = Value[0:4] + Value[5:7] + Value[8:10] + Value[11:13] + Value[14:16] + Value[17:19]
if any(Character not in "0123456789" for Character in NumberText):
raise ValueError("time needs ordinary digits")
Year = Digits(Value[0:4], 9999)
Month = Digits(Value[5:7], 12)
Day = Digits(Value[8:10], 31)
Hour = Digits(Value[11:13], 23)
Minute = Digits(Value[14:16], 59)
Second = Digits(Value[17:19], 59)
try:
DateTime(Year, Month, Day)
except ValueError:
raise ValueError("date does not exist") from None
try:
DateTime(Year, Month, Day, Hour, Minute, Second)
except ValueError:
raise ValueError("time does not exist") from None
def ParseEntry(Entry):
try:
if len(Entry.Original) > 1024:
raise ValueError("line exceeds 1024 bytes")
if any(Character != "\t" and not 32 <= ord(Character) <= 126
for Character in Entry.Original):
raise ValueError("unsupported byte in practice format")
Fields = Entry.Original.split("\t")
if len(Fields) != 4:
raise ValueError("expected exactly four tab-separated fields")
CheckTime(Fields[0])
CheckAddress(Fields[1])
CheckAddress(Fields[2])
Entry.Port = Digits(Fields[3], 65535)
if Entry.Port == 0:
raise ValueError("port must be 1 to 65535 in this format")
Entry.EventTime, Entry.Source, Entry.Destination = Fields[:3]
Entry.Accepted = True
except ValueError as Error:
Entry.Problem = str(Error)
return Entry
def ReadLog(FileName):
with Path(FileName).open("rb") as Input:
Data = Input.read(1048577)
if len(Data) > 1048576:
raise ValueError("input exceeds 1 MiB teaching limit")
if not Data:
raise ValueError("empty input")
Lines = Data.split(b"\n")
if Lines[-1] == b"":
Lines.pop()
Lines = [Line[:-1] if Line.endswith(b"\r") else Line for Line in Lines]
if Lines[0] != b"Time\tSource\tDestination\tPort":
raise ValueError("unsupported header")
if len(Lines) - 1 > 1000:
raise ValueError("more than 1000 data lines")
return [ParseEntry(LogEntry(Number, Line.decode("latin-1")))
for Number, Line in enumerate(Lines[1:], start=2)]
Choose what one group means
One group is the exact combination of source address, destination address and destination port. Two events with the same source but different destinations belong to different groups. Two events to the same destination but different ports also stay separate.
This matters because the output is a count of entries. The practice file has no protocol, source port, byte count, session identifier or unique event identifier. It cannot tell whether two entries describe separate sessions, retries, duplicated logging or something else. It does not infer a service merely from a familiar port number.
Before adding another field to the grouping rule, decide which question you want the report to answer. A summary by source alone is a different report. Calling both reports "top connections" without explaining the difference makes the output easy to misread.
Write the summary program
Save this as SummariseLog.py beside LogReader.py:
import sys as Sys
from LogReader import ReadLog
def Main():
if len(Sys.argv) != 2:
print("Usage: python3 SummariseLog.py practice.tsv", file=Sys.stderr)
return 2
try:
Entries = ReadLog(Sys.argv[1])
except (OSError, ValueError) as Error:
print(f"Input stopped: {Error}", file=Sys.stderr)
return 2
Groups = {}
Good = 0
Bad = 0
for Entry in Entries:
if not Entry.Accepted:
Bad += 1
print(f"Rejected line {Entry.Number}: {Entry.Problem}")
else:
Good += 1
Key = (Entry.Source, Entry.Destination, Entry.Port)
Groups[Key] = Groups.get(Key, 0) + 1
print("Count Source Destination Port")
for Key, Count in sorted(Groups.items(), key=lambda Item: -Item[1]):
Source, Destination, Port = Key
print(f"{Count} {Source} {Destination} {Port}")
print(f"Accepted: {Good}; rejected: {Bad}. Counts are log entries, not sessions.")
return 1 if Bad else 0
if __name__ == "__main__":
raise SystemExit(Main())
The code has no comments; the article explains its work. For each accepted entry, it looks up that source, destination and port group, starting from zero if it has not been seen yet, then adds one. Rejected entries increment a separate rejected total and print their physical line numbers and reasons.
The final sorted call orders groups by descending count. Python preserves insertion order in this dictionary, and its sort keeps equal-count items in that order. Tied groups therefore appear in the order they were first accepted from the input, not in an order of threat or importance.
Run the program
Install Python 3 through the supported route for your computer. With both source files in one folder:
python3 SummariseLog.py practice.tsv
No extra packages are needed. The actual runs used Python 3.10.12 on Linux. Windows and macOS were not tested. Example names use PascalCase, such as Groups and Main; Python's own built-in names, library methods and special __name__ spelling stay unchanged.
Actual output:
Rejected line 5: date does not exist
Rejected line 6: number outside supported range
Count Source Destination Port
2 192.0.2.10 198.51.100.20 443
1 192.0.2.11 203.0.113.30 53
Accepted: 3; rejected: 2. Counts are log entries, not sessions.
Three accepted entries become two groups. Two malformed entries do not disappear: their reasons appear before the table and their count stays in the summary. The program returns 1 because the input is partly rejected. Fully accepted input returns 0, and an overall load failure returns 2 with no table.
The accepted counts sum to 3, the accepted-entry total. That is one useful consistency check. It does not reconcile the log against a network sensor or establish that the file includes every relevant event.
Check the ordering with a second invented file
A separate test file had one event in the first group, three in the second and two in the third. Actual output:
Count Source Destination Port
3 192.0.2.2 198.51.100.2 80
2 192.0.2.3 198.51.100.3 53
1 192.0.2.1 198.51.100.1 443
Accepted: 6; rejected: 0. Counts are log entries, not sessions.
The descending 3, 2, 1 order was checked automatically. The rebuilt Python reader and summary also passed 72 shared command-line checks covering valid/invalid format cases, limits and missing input. Exact repeated lines were counted twice, not silently removed. Input bytes stayed unchanged.
Keeping repetitions is intentional. Without a reliable unique event reference, a repeated line may be duplicated logging or a genuinely repeated action. Decide how to handle that separately, with evidence; do not make a summary look cleaner by assuming the answer.
Investigate the result without making a threat verdict
A busy group may be ordinary business activity. A quiet group may still deserve attention. A count alone does not establish intent, data volume, unauthorised access or a scan. Compare with the actual logging format, the period covered and the context of the systems you are allowed to inspect.
This program groups the whole practice file. It does not compute a rate because it has no chosen observation interval. Ten entries over a minute and ten over a day are different questions. The next planned guide adds explicit time-window work and tests its boundaries rather than simply calling a large count a burst.
Missing evidence changes the meaning too. If many rows were rejected, the table is a summary of what the parser accepted, not the full file. Read the reasons, correct the source/export through an authorised process if appropriate, and rerun. Preserve the original rather than editing it until the rejects vanish.
Keep the exercise small and read-only
The shared reader reads at most 1 MiB plus one byte, rejecting input beyond its 1 MiB teaching limit. It accepts at most 1,000 data lines and rejects a data line longer than 1,024 bytes after reading. It loads entries and groups in memory. A file changing while it is read is not a protected snapshot. Use a stable invented input; this is not a hardened parser for hostile growing logs or a tested large-file tool.
The programs themselves only read the input. A command-window redirection can still overwrite it before they start. If you save output, choose a new report filename, never the input name. Real logs and output can reveal sensitive system activity; keep them in authorised protected storage.
This guide's useful outcome is a reproducible grouping rule with its gaps visible.
References
- Python pathlib: opening files.
- Python dataclasses: named record fields.
- RFC 5737: documentation-only address ranges.
- RFC 3339: timestamp reference. The reader handles only the exact UTC spelling described here.
No claim of real export compatibility, network completeness, malicious activity, session counts or production performance is made.
Top comments (0)