This is a submission for the Kaggle Benchmarking Challenge.
What I Benchmarked
Can a language model use the relationships around a transaction to help identify fraud?
I built a benchmark using transaction data from the IEEE-CIS Fraud Detection dataset. Each example presents a target transaction alongside a small text representation of its local network neighborhood. The neighboring transactions are connected through shared attributes such as card, address, or device information. The model does not receive the ground-truth labels in its prompt.
The benchmark measures two behaviors: classifying the target transaction as fraudulent or legitimate, and identifying a potentially suspicious neighboring transaction. I was interested in whether a general-purpose language model could make useful predictions from graph context after that context had been translated into text.
Models Tested
I ran Gemma 4 31B (google/gemma-4-31b) on 100 samples, with both tasks evaluated for each sample. I chose it as the benchmark’s target language model to test whether a large general-purpose model could interpret transaction relationships from the text prompt.
I also compared its classification results with the project’s GraphSAGE graph-neural-network baseline. This was a single-model evaluation, not a broad ranking across language models.
Findings
The run completed with 200 non-empty responses: 100 for transaction classification and 100 for suspicious-neighbor identification. A controlled probe returned text with no token override but an empty response when the 128-token cap was nested under extra_body. The completed run sent the cap as a top-level max_tokens parameter and produced all 200 responses.
Gemma achieved 55% classification accuracy and 53% accuracy on the fraud-neighbor proxy. The 100 classification labels were evenly split—50 fraudulent and 50 legitimate—so the majority-class baseline was 50%. At 55 correct out of 100, this result does not establish performance above that reference. No baseline comparison is claimed for the fraud-neighbor proxy.
On classification, the GraphSAGE baseline was correct on 67 of the 100 samples. Both systems were correct on 44; Gemma was correct when GraphSAGE was wrong on 11; and GraphSAGE was correct when Gemma was wrong on 23. Both missed 22. In this sample, the graph-based baseline outperformed Gemma.
The neighbor score needs careful interpretation. IEEE-CIS does not provide verified fraud-ring membership or ring-leader labels, so I scored the selected neighbor using its transaction fraud label. That makes 53% a fraud-neighbor proxy score, not evidence that the model can identify actual fraud rings or their leaders.
The main takeaway is that providing a local transaction neighborhood as text did not, by itself, produce strong results from this model on this sample. I would next test more models, increase the sample size, and run prompt ablations to measure how much the neighborhood details affect predictions.
Top comments (0)