DEV Community

Deep Learning

This tag is for discussing, sharing articles, and asking questions primarily on deep learning - a subfield of machine learning.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Transformers: Understanding the Architecture Behind Modern AI

Transformers: Understanding the Architecture Behind Modern AI

Comments
7 min read
Deep Learning from Scratch in 1400 Lines - Neve & Frost Framework

Deep Learning from Scratch in 1400 Lines - Neve & Frost Framework

Comments 1
10 min read
What optimizer.step() Actually Does: One Step of SGD, Momentum and Adam by Hand

What optimizer.step() Actually Does: One Step of SGD, Momentum and Adam by Hand

Comments 1
10 min read
MULTIPITA: Reorganizing Compute Without Compressing Identity

MULTIPITA: Reorganizing Compute Without Compressing Identity

1
Comments
1 min read
Incin, a rust machine learning framework for setting fire to dimensionality bugs

Incin, a rust machine learning framework for setting fire to dimensionality bugs

Comments
1 min read
Entropy and Cross-Entropy, Explained

Entropy and Cross-Entropy, Explained

Comments
3 min read
LTX-2.5: How a Diffusion-Based Video Decoder Changes the Open-Weights Video Generation Stack

LTX-2.5: How a Diffusion-Based Video Decoder Changes the Open-Weights Video Generation Stack

Comments
5 min read
The AI Revolution Isn’t Coming — It’s Already Here. Are You Ready to Build With It?

The AI Revolution Isn’t Coming — It’s Already Here. Are You Ready to Build With It?

Comments
3 min read
On-Policy vs Off-Policy Training of LLMs: How Models Start Learning From Their Own Outputs

On-Policy vs Off-Policy Training of LLMs: How Models Start Learning From Their Own Outputs

1
Comments 1
11 min read
How to Generate Synthetic Time-Series Data Without Faking the Signal

How to Generate Synthetic Time-Series Data Without Faking the Signal

Comments
8 min read
QUASAR: How Saliency-Weighted Reconstruction Closes the Loss Floor Gap in LLM Quantization-Aware Training

QUASAR: How Saliency-Weighted Reconstruction Closes the Loss Floor Gap in LLM Quantization-Aware Training

Comments
5 min read
how i built a graph neural network (and what i learned)

how i built a graph neural network (and what i learned)

Comments
4 min read
Transformer Architecture Basics

Transformer Architecture Basics

Comments
8 min read
We measured sharpness during loss spikes. The instrument started returning negative numbers.

We measured sharpness during loss spikes. The instrument started returning negative numbers.

Comments
3 min read
Masked Self-Attention, Explained Through Avengers: Endgame

Masked Self-Attention, Explained Through Avengers: Endgame

1
Comments
6 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.