DEV Community

Lawson Dong
Lawson Dong

Posted on AI-assisted

Training My First Neural Network on Windows with WSL 2 and PyTorch

As a physics undergraduate beginning to explore AI research, I wanted to understand the practical workflow behind a neural network experiment: where the code lives, how the environment works, how to run training, and what gets saved afterward.

This tutorial brings together my setup notes and first PyTorch experiment. We will build a small network for XOR, starting from a Windows computer and ending with a saved model file. The focus is on getting the workflow running; the mathematics can come later.

You need a supported Windows 10 or Windows 11 installation, permission to install WSL, and an internet connection. A GPU is unnecessary for this four-example experiment.

1. Understand the tools before installing them

Overview of the local Windows-to-PyTorch training workflow

The workflow is Windows terminal → WSL 2 → Ubuntu → project folder → Python virtual environment → PyTorch → training results.

Tool Role in this experiment
CMD / PowerShell Launches WSL from Windows
WSL 2 Runs a Linux kernel using lightweight virtualization
Ubuntu Provides the Linux operating environment
APT Installs Ubuntu system packages
Python Executes our training script
venv Isolates the project's Python dependencies
pip Installs packages inside that environment
PyTorch Provides tensors, neural network layers, and automatic differentiation

Docker is an optional next step for packaging environments. Ordinary Ubuntu and PyTorch development works without Docker Desktop running.

2. Install and launch Ubuntu

Run these commands in Windows CMD or PowerShell:

wsl -l -v
Enter fullscreen mode Exit fullscreen mode

On my machine, the only listed distribution was initially docker-desktop. That did not mean I already had an Ubuntu development environment.

To install Ubuntu, open PowerShell as administrator and run:

wsl --install -d Ubuntu
Enter fullscreen mode Exit fullscreen mode

Restart Windows if prompted. On Ubuntu's first launch, create a Linux username and password. The password input displays no characters or asterisks.

Check the installation from Windows:

wsl -l -v
Enter fullscreen mode Exit fullscreen mode

Confirm that Ubuntu is listed with VERSION 2. If it uses version 1, run:

wsl --set-version Ubuntu 2
Enter fullscreen mode Exit fullscreen mode

Optionally make Ubuntu the default, then launch it:

wsl --set-default Ubuntu
wsl -d Ubuntu
Enter fullscreen mode Exit fullscreen mode

For installation requirements and troubleshooting, see Microsoft's WSL installation guide.

3. Find your Linux home directory

From this point onward, commands run inside Ubuntu, unless explicitly labeled otherwise.

A terminal prompt might look like:

lawson@computer:~$
Enter fullscreen mode Exit fullscreen mode

lawson is the Linux username, computer is the hostname, and ~ means the user's home directory. Do not copy the prompt itself when running commands.

Linux file system structure, including the user's home directory

Your Linux home directory is typically /home/<username>. Windows drives are accessible through paths such as /mnt/c/. For this experiment, keep the project in your Linux home directory.

cd ~
pwd
ls
Enter fullscreen mode Exit fullscreen mode
Command What it does
pwd Prints the current working directory
ls Lists directory contents
cd Changes directories
mkdir Creates directories
touch Creates an empty file or updates its timestamps
cat Displays or combines file contents
sudo Runs a command as another user, usually root

4. Install system tools with APT

APT installs system packages from Ubuntu software repositories

APT manages Ubuntu packages. Refresh its package information and install the tools we need:

sudo apt update
sudo apt install -y git python3 python3-pip python3-venv nano
Enter fullscreen mode Exit fullscreen mode

Then check:

python3 --version
git --version
Enter fullscreen mode Exit fullscreen mode

APT handles system software; pip handles Python packages. In the following steps, pip installs packages into our project environment.

5. Create a project and virtual environment

cd ~
mkdir -p research/first-neural-network
cd research/first-neural-network
pwd
Enter fullscreen mode Exit fullscreen mode

The path should end with research/first-neural-network.

Separate virtual environments let projects use different PyTorch versions

The diagram's version numbers are examples: each project can manage its own dependencies independently.

Create and activate the environment:

python3 -m venv .venv
source .venv/bin/activate
Enter fullscreen mode Exit fullscreen mode

Your prompt should now begin with (.venv). Closing the terminal leaves this directory on disk; you will reactivate it in the next session.

6. Install and verify PyTorch

With .venv active, install a CPU build of PyTorch and NumPy:

python -m pip install --upgrade pip
python -m pip install torch --index-url https://download.pytorch.org/whl/cpu
python -m pip install numpy
Enter fullscreen mode Exit fullscreen mode

Check PyTorch's official installation selector for supported Python versions or a GPU-specific installation command.

Verify the installation:

python -c "import torch, numpy; print('PyTorch:', torch.__version__); print('NumPy:', numpy.__version__)"
python -c "import torch; print(torch.rand(2, 2))"
Enter fullscreen mode Exit fullscreen mode

The second command should print a random 2 × 2 tensor. In my original setup, tensor creation worked but PyTorch warned that NumPy was missing. Installing NumPy resolved that missing dependency.

7. Write the training script

Our task is XOR: output 1 when the two inputs differ and 0 when they match.

Input Target
[0, 0] 0
[0, 1] 1
[1, 0] 1
[1, 1] 0

Open a new file:

nano train.py
Enter fullscreen mode Exit fullscreen mode

Paste this code:

import torch
import torch.nn as nn

torch.manual_seed(42)

# 1. Training data
X = torch.tensor([
    [0., 0.],
    [0., 1.],
    [1., 0.],
    [1., 1.]
])
y = torch.tensor([[0.], [1.], [1.], [0.]])

# 2. A small network: two inputs, four hidden units, one output
model = nn.Sequential(
    nn.Linear(2, 4),
    nn.Tanh(),
    nn.Linear(4, 1)
)

# 3. Loss and optimizer
criterion = nn.BCEWithLogitsLoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.05)

# 4. Training
model.train()
for epoch in range(2001):
    optimizer.zero_grad()
    logits = model(X)
    loss = criterion(logits, y)
    loss.backward()
    optimizer.step()

    if epoch % 200 == 0:
        print(f"Epoch {epoch}, Loss: {loss.item():.6f}")

# 5. Inspect predictions on the four training examples
model.eval()
with torch.no_grad():
    probabilities = torch.sigmoid(model(X))
    predicted_labels = (probabilities >= 0.5).int()
    print("\nProbabilities:")
    print(probabilities)
    print("\nPredicted labels:")
    print(predicted_labels)

# 6. Save learned parameters
torch.save(model.state_dict(), "xor_model.pth")
print("\nSaved weights to xor_model.pth")
Enter fullscreen mode Exit fullscreen mode

Save with Ctrl + O, press Enter, and exit with Ctrl + X.

The loop computes predictions, measures the error, computes gradients, and updates parameters. BCEWithLogitsLoss takes the network's raw output, so the sigmoid is applied when inspecting probabilities afterward.

This is a small network with one hidden layer. It is a first neural network training exercise, rather than a large-scale deep learning experiment.

8. Run training and inspect the result

python train.py
Enter fullscreen mode Exit fullscreen mode

One mistake I made was trying torch train.py. PyTorch is a Python library; Python executes the file.

You should see periodic loss reports, followed by probabilities and labels. Check that the loss trends downward and the predicted labels are:

tensor([[0],
        [1],
        [1],
        [0]], dtype=torch.int32)
Enter fullscreen mode Exit fullscreen mode

Exact probabilities and loss values can vary with software versions. These predictions use the same four examples used for training, so they check that the network learned the XOR truth table; they do not measure generalization to a separate dataset.

Check the files:

ls -lh train.py xor_model.pth
Enter fullscreen mode Exit fullscreen mode

xor_model.pth contains the model's parameter state dictionary. It does not contain the training script or a complete experiment checkpoint. To restore these weights later, recreate the same architecture and load the state dictionary; see PyTorch's saving and loading tutorial.

9. Resume the project next time

In Windows CMD or PowerShell:

wsl -d Ubuntu
Enter fullscreen mode Exit fullscreen mode

Inside Ubuntu:

cd ~/research/first-neural-network
source .venv/bin/activate
python train.py
Enter fullscreen mode Exit fullscreen mode

This runs training again from newly initialized weights. It does not resume the previously saved model automatically.

To leave the environment, run deactivate. To leave Ubuntu, run exit. An active foreground process may terminate if its terminal closes, but files already saved remain on disk.

10. Common setup mistakes

Problem Fix
torch: command not found Run python train.py
No module named 'torch' Activate .venv, then install PyTorch with that environment's Python
No module named 'numpy' Run python -m pip install numpy in .venv
train.py prints nothing Check that the file contains and saves the code
wsl -l -v fails inside Ubuntu Run it from Windows, or use wsl.exe -l -v inside Ubuntu
Docker Desktop is stopped Docker is optional for this experiment
The terminal was closed Reopen Ubuntu, return to the folder, and reactivate .venv

11. Where this fits in a research workflow

Broader training workflow with Windows, WSL 2, Ubuntu, Docker, and a GPU

This diagram from my notes shows a possible later setup with Docker and GPU acceleration. Those components are optional extensions beyond the CPU workflow used here.

Research workflow connecting version control, local development, remote compute, and saved results

For larger experiments, I want to connect local development and version control with remote compute, then keep model weights, metrics, figures, and logs together. Access to university computing platforms depends on their own eligibility and allocation rules.

The useful milestone here is modest but concrete: a project directory, isolated dependencies, an executable training script, inspectable predictions, and saved parameters.

My notes and projects: GitHub · Personal website

Top comments (0)