DEV Community

Machine Learning

A branch of artificial intelligence (AI) and computer science which focuses on the use of data and algorithms to imitate the way that humans learn, gradually improving its accuracy.

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Plain Gemma 4 26B vs Jev on One EC2 L4: 2.1 Points Behind Overall, Level on Yes/No, 4.5 Behind on Multiple Choice

Plain Gemma 4 26B vs Jev on One EC2 L4: 2.1 Points Behind Overall, Level on Yes/No, 4.5 Behind on Multiple Choice

1
Comments
23 min read
Plain Gemma 4 26B vs Jev on One EC2 L4: 2.1 Points Behind Overall, Level on Yes/No, 4.5 Behind on Multiple Choice

Plain Gemma 4 26B vs Jev on One EC2 L4: 2.1 Points Behind Overall, Level on Yes/No, 4.5 Behind on Multiple Choice

1
Comments
23 min read
25 LLM architecture blocks, side by side, in runnable PyTorch

25 LLM architecture blocks, side by side, in runnable PyTorch

Comments
8 min read
Detección de fraude: por qué el AUC-ROC engaña y el AUC-PR no (datos reales)

Detección de fraude: por qué el AUC-ROC engaña y el AUC-PR no (datos reales)

Comments
2 min read
In 2018 I hand-wrote a C++ deep learning framework so I'd never pad a batch. In 2023 LLM serving landed on the same structure.

In 2018 I hand-wrote a C++ deep learning framework so I'd never pad a batch. In 2023 LLM serving landed on the same structure.

Comments
13 min read
I didn't fix the bug: contributing to a 20k-star ML repo by measuring it

I didn't fix the bug: contributing to a 20k-star ML repo by measuring it

Comments
11 min read
Jev vs LLMs: Why AI Agents May Need a Decision Layer

Jev vs LLMs: Why AI Agents May Need a Decision Layer

4
Comments
5 min read
Can an AI Agent Know When Not to Act? A Fail-Closed Reliability Benchmark Across Six Models

Kaggle Benchmarking Challenge Submission

Can an AI Agent Know When Not to Act? A Fail-Closed Reliability Benchmark Across Six Models

Comments
4 min read
Open‑vocab relation predictor triples recall

Open‑vocab relation predictor triples recall

Comments
2 min read
Finding fraud rings a risk model can't see: an agentic investigator on TigerGraph

Finding fraud rings a risk model can't see: an agentic investigator on TigerGraph

Comments
6 min read
Muse ของ Meta เอเจนต์ส่วนตัวที่ทำงานแทนคุณจริง ตั้งแต่ส่งอีเมลถึงต่อรอง

Muse ของ Meta เอเจนต์ส่วนตัวที่ทำงานแทนคุณจริง ตั้งแต่ส่งอีเมลถึงต่อรอง

Comments
2 min read
Three perfect scores weren't enough: testing AI outage decisions one fact at a time

Kaggle Benchmarking Challenge Submission

Three perfect scores weren't enough: testing AI outage decisions one fact at a time

Comments
5 min read
SEO อัตโนมัติด้วย AI Agent ปี 2026 ทำได้จริงถึงไหน

SEO อัตโนมัติด้วย AI Agent ปี 2026 ทำได้จริงถึงไหน

Comments
1 min read
Abliterated ชุมชนปลดล็อกโมเดล Qwen3.8 เอง เทคนิคและคำถามที่ตามมา

Abliterated ชุมชนปลดล็อกโมเดล Qwen3.8 เอง เทคนิคและคำถามที่ตามมา

Comments
1 min read
Top 7 Auto Routing Platforms for Model Selection in 2026

Top 7 Auto Routing Platforms for Model Selection in 2026

Comments
15 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.