This is a submission for the Kaggle Benchmarking Challenge.
What I Benchmarked
Can a language model use the relationships around a transaction to help identify fraud?
I built a benchmark using transaction data from the IEEE-CIS Fraud Detection dataset. Each example presents a target transaction alongside a small text representation of its local network neighborhood. The neighboring transactions are connected through shared attributes such as card, address, or device information. The model does not receive the ground-truth labels in its prompt.
The benchmark measures two behaviors: classifying the target transaction as fraudulent or legitimate, and identifying a potentially suspicious neighboring transaction. I was interested in whether a general-purpose language model could make useful predictions from graph context after that context had been translated into text.
Models Tested
I ran Gemma 4 31B (google/gemma-4-31b) on 100 samples, with both tasks evaluated for each sample. I chose it as the benchmark’s target language model to test whether a large general-purpose model could interpret transaction relationships from the text prompt.
I also compared its classification results with the project’s GraphSAGE graph-neural-network baseline. This was a single-model evaluation, not a broad ranking across language models.
Findings
The run completed with 200 non-empty responses: 100 for transaction classification and 100 for suspicious-neighbor identification. A controlled probe returned text with no token override but an empty response when the 128-token cap was nested under extra_body. The completed run sent the cap as a top-level max_tokens parameter and produced all 200 responses.
Gemma achieved 55% classification accuracy and 53% accuracy on the fraud-neighbor proxy. The 100 classification labels were evenly split—50 fraudulent and 50 legitimate—so the majority-class baseline was 50%. At 55 correct out of 100, this result does not establish performance above that reference. No baseline comparison is claimed for the fraud-neighbor proxy.
On classification, the GraphSAGE baseline was correct on 67 of the 100 samples. Both systems were correct on 44; Gemma was correct when GraphSAGE was wrong on 11; and GraphSAGE was correct when Gemma was wrong on 23. Both missed 22. In this sample, the graph-based baseline outperformed Gemma.
The neighbor score needs careful interpretation. IEEE-CIS does not provide verified fraud-ring membership or ring-leader labels, so I scored the selected neighbor using its transaction fraud label. That makes 53% a fraud-neighbor proxy score, not evidence that the model can identify actual fraud rings or their leaders.
The main takeaway is that providing a local transaction neighborhood as text did not, by itself, produce strong results from this model on this sample. I would next test more models, increase the sample size, and run prompt ablations to measure how much the neighborhood details affect predictions.
Top comments (1)
Official Platform Update
Security protocols have been updated for all developer accounts.