This is my submission for the Hacktoberfest 2026 Weekend Challenge: Build for a Friend.
What I Built
My family gets scam messages all the time.
Fake bank alerts. KYC requests. Delivery messages asking for customs fees. Messages from unknown numbers pretending to be someone we know. Even fake police or cybercrime notices threatening arrest unless money is paid immediately.
Most people in my family aren't technical. They aren't going to inspect a URL or check whether a domain actually belongs to a bank.
Sometimes they just ask me:
"Is this real?"
So I built something for that exact situation.
ScamLens is a scam and social engineering analyzer for messages and screenshots. You can paste a suspicious message or upload a screenshot, and it analyzes the content locally and explains why it thinks the message is suspicious or legitimate.
The important part is that it doesn't require sending the message to a cloud AI service.
The analysis runs on the user's own machine.
Demo
Here is the demo:
The demo covers three situations:
- A phishing message pretending to be from a bank.
- A social engineering scam without a suspicious link.
- A legitimate OTP security notification.
The third case ended up being the most important one.
The part that changed the project
I thought I was basically done after getting the main detection system working.
I had a small 16 case evaluation set with 8 scam messages and 8 legitimate messages. After calibration, all 16 were classified correctly.
Then I handed the laptop to a non technical family member and let them try the live application.
I didn't explain what the messages were supposed to be.
The first phishing message looked convincing.
The second social engineering message also looked believable.
Then came the legitimate OTP.
ScamLens flagged it as:
SCAM DETECTED
The message said:
Your login OTP is 482913. Valid for 5 minutes. DO NOT share this OTP with anyone, including bank employees or support agents.
There was no suspicious URL. No payment request. No request to send the OTP to anyone.
The family member immediately questioned the result. If the message was telling her not to share the OTP, why was ScamLens treating that instruction as something a scammer was demanding?
She was right.
When I looked at the result more closely, the problem became obvious.
The model had interpreted:
- "Valid for 5 minutes" as urgency.
- "DO NOT share this OTP" as social engineering.
- The security warning as if it were a request from the sender.
Even worse, the deterministic security checks had found no suspicious signals.
The system was basically saying:
There is no technical evidence of a scam, but I'm calling it a scam anyway.
That was a real bug.
A security tool that constantly warns people about legitimate messages can become a problem of its own. Eventually, people stop trusting the warnings.
So I went back into the reasoning layer.
The fix wasn't just to add a rule saying "DO NOT share" means safe. That would be too easy to bypass.
Instead, I added explicit semantic fields to distinguish things like:
- Is the sender asking the user to reveal a secret?
- Is the message warning the user not to reveal a secret?
- Is the OTP expiry simply describing how long a code is valid?
- Is the urgency actually coercive?
- Is the security advice protective?
The decision engine then uses these signals together instead of treating individual words as proof of a scam.
After the fix, the same OTP message returned:
LIKELY LEGITIMATE
The phishing and social engineering examples still returned:
SCAM DETECTED
That was probably the most useful test of the entire project.
A real person found a problem that my carefully prepared test set hadn't.
How I Built It
I didn't want the language model to be the final authority.
For a text message, ScamLens uses Llama 3 locally through Ollama, then runs deterministic security checks and combines the results through a decision engine.
For screenshots, EasyOCR first extracts the text and then the same analysis takes place.
The model looks for things that are difficult to capture with simple rules:
- Authority impersonation
- Social engineering
- Emotional manipulation
- Fear and panic tactics
- Legal intimidation
- Financial coercion
- Suspicious requests
The deterministic verifier checks concrete security signals such as:
- Suspicious URLs
- Unencrypted HTTP links
- Suspicious domains
- Credential requests
- OTP and PIN harvesting
- Payment requests
- URL shorteners
- Brand and domain mismatches
The final verdict comes from combining both.
The main idea is simple:
The model understands the message. The verifier checks the message. The decision engine combines the evidence.
The stack is:
- Llama 3 8B
- Ollama
- EasyOCR
- FastAPI
- React
- TypeScript
- Vite
One surprisingly important part was OCR.
A screenshot can look perfectly readable to a person while still being difficult for OCR. During testing, a URL such as http://sbi-kyc.xyz/verify was initially read incorrectly enough that URL detection failed.
I added image preprocessing before OCR, including upscaling, grayscale and contrast processing, and sharpening.
That made the screenshot analysis more reliable.
Why Does Open Innovation Matter?
This project deals with information I wouldn't want to casually send to another service.
A suspicious message might contain:
- An OTP
- A phone number
- Account information
- Transaction details
- Someone's name
- Private conversations
If someone wants to ask, "Is this message a scam?", I don't think sending the entire message to a third party AI service should have to be the default answer.
That's why local open source AI mattered for ScamLens.
Llama 3 can run locally through Ollama, so the language model analysis can happen on the user's own machine without requiring a cloud AI API.
It also gave me the flexibility to build the rest of the system around the model.
I didn't have to treat the model's answer as the final truth.
I could combine it with deterministic security checks, structured outputs, and a decision layer that I could test and change myself.
That became especially important because the model was capable of detecting things that simple rules could not, but it also had its own blind spots.
The project needed engineering around the model.
What Actually Works
I tested ScamLens against a small curated evaluation set with:
- 8 scam cases
- 8 legitimate cases
- 16 cases in total
This is a calibration and evaluation set, not an industry benchmark.
After calibration:
| Approach | Scam Recall | Legitimate False Positives |
|---|---|---|
| Deterministic verifier only | 75% (6/8) | 0% (0/8) |
| Local Llama 3 only | 100% (8/8) | 62.5% (5/8) |
| Combined decision engine | 100% (8/8) | 0% (0/8) |
The interesting result wasn't simply that the combined system performed better.
The two layers caught different things.
The deterministic verifier could catch obvious technical signals very well.
But the pure social engineering case had no suspicious URL and no payment keyword. The rule based system alone classified it as legitimate.
The local AI caught the family impersonation and credential extraction pattern.
So the AI isn't there just because it sounds impressive.
It catches a category of manipulation that simple security rules can miss.
The current release has:
- 51/51 backend tests passing
- 5/5 frontend tests passing
- Production build passing
- Technical phishing classified as SCAM DETECTED
- Social engineering scam classified as SCAM DETECTED
- Legitimate OTP test classified as LIKELY LEGITIMATE
- Screenshot analysis through EasyOCR
- Local Llama 3 inference through Ollama
What I Learned
The biggest thing I learned is that curated tests and real users find different bugs.
My 16 case test set was designed to cover different scam patterns and legitimate controls.
It passed.
Then a family member used the actual interface and broke it on the third message.
The interesting part wasn't just that the model was wrong.
The output itself was confusing.
A non technical person doesn't look at a verdict the same way the developer who built the system does. They are trying to answer a much simpler question:
"Should I trust this message?"
That changes how the entire interface and reasoning system needs to be designed.
I also learned that negative security advice is surprisingly difficult for a language model to interpret correctly.
"Send me your OTP" and "DO NOT share your OTP" contain many of the same words.
But one is an attack and the other is protection.
The difference isn't in the vocabulary.
It's in the intent.
That distinction ended up being one of the most important parts of ScamLens.
Limitations
ScamLens is an aid, not a guarantee.
There are still several limitations:
- The evaluation set is only 16 curated cases, not a production benchmark.
- Local Llama 3 inference takes roughly 26 seconds per text message on my hardware.
- Screenshot analysis can take longer because of OCR and preprocessing.
- OCR can struggle with heavily distorted images, unusual fonts, or low contrast screenshots.
- Sophisticated scams can deliberately imitate legitimate security language.
- The current system analyzes one message at a time and does not understand a long multi turn conversation.
- Running the project locally requires installing Ollama and the model manually.
No scam detector can guarantee that a message is safe.
The goal of ScamLens is to give a person useful evidence and a clearer explanation before they act, not to replace their judgment completely.
Code & Demo
GitHub:
https://github.com/Ritinpaul/ScamLens
Demo:
https://youtu.be/EiPYk873mYY
ScamLens started with a simple question:
What if someone in my family receives a message that looks completely real, and they don't know whether to trust it?
The answer wasn't just to build a scam detector.
It was to build something they could actually use, keep the analysis private, explain the reasoning, and listen when a real person showed me that the system was getting something wrong.
That's what Build for a Friend meant to me.
Top comments (0)