DEV Community

Nishant Gaurav
Nishant Gaurav

Posted on AI-assisted

A Beginner to Intermediate Guide to Unit Testing in Python with Pytest

If you write Python code, at some point you will need a reliable way to check that it actually works, not just once by hand, but every time you change something. That is what unit testing gives you, and pytest is one of the most popular tools for writing those tests in Python. This guide walks through everything from installing pytest and writing your very first test, through fixtures, parameterized tests, and mocking, including mocking API calls, since those are patterns you will genuinely use once you start testing real projects.


Why Pytest Instead of Python's Built In unittest

Python actually ships with a built in testing module called unittest. It works, and plenty of projects use it. But its syntax is more verbose and less intuitive than pytest's. You typically have to write test classes that inherit from unittest.TestCase, and assertions use method calls like self.assertEqual(a, b) instead of a plain assert statement.

Pytest is easier to pick up, requires less boilerplate, and has become the standard choice for testing in the Python ecosystem. The core testing ideas covered in this guide (writing isolated tests, using fixtures, mocking dependencies) apply no matter which framework you end up using. Pytest just makes the syntax for expressing those ideas simpler.


Setting Up Pytest

Before writing any tests, you need Python installed, and then you install pytest itself as a package, just like any other Python library.

Open a terminal or command prompt and run:

pip install pytest
Enter fullscreen mode Exit fullscreen mode

On Mac or Linux, depending on your setup, you might need to use pip3 instead:

pip3 install pytest
Enter fullscreen mode Exit fullscreen mode

Alongside pytest, this guide also uses a companion package called pytest mock, which adds mocking support, covered in detail later. Install it the same way:

pip install pytest-mock
Enter fullscreen mode Exit fullscreen mode

or on Mac or Linux:

pip3 install pytest-mock
Enter fullscreen mode Exit fullscreen mode

With both installed, you are ready to write and run tests.

Writing Your First Test

Say you have a project folder, call it python_testing, opened in your code editor. Inside it, you create a file called main.py, which contains the actual code you want to test.

# main.py

def get_weather(temp):
    if temp > 20:
        return "hot"
    else:
        return "cold"
Enter fullscreen mode Exit fullscreen mode

This is a small function, but it is exactly the kind of thing worth testing: given an input, it should reliably produce the expected output.

The Test File Naming Convention

To test this function, you create a second file. Pytest has a naming convention it looks for automatically: a test file should be named test_ followed by the name of the module it is testing. Since the code lives in main.py, the test file is named:

test_main.py
Enter fullscreen mode Exit fullscreen mode

This prefix matters. It is literally how pytest discovers which files contain tests when you run it. Generally you will have one test file per module or file you want to test, though you can also organize tests around individual functions if a module grows large.

Writing the Test Itself

Inside test_main.py, there are really only two things you need to do to write a working test: import the code you want to test, and write a function containing an assertion.

# test_main.py

from main import get_weather

def test_get_weather():
    assert get_weather(21) == "hot"
Enter fullscreen mode Exit fullscreen mode

Breaking this down:

  • from main import get_weather imports the function being tested.
  • def test_get_weather(): defines the test function. Like the file itself, the function name is prefixed with test_. This is what tells pytest, this function is a test, run it.
  • assert get_weather(21) == "hot" is the actual check. It calls get_weather with 21, and asserts that the result equals "hot".

What assert Actually Does

An assertion checks whether a condition is True or False. If the condition evaluates to True, the test case passes silently. If it evaluates to False, the test fails, and pytest reports exactly what went wrong.

assert <condition>

condition is True  -> test passes
condition is False -> test fails, with details about what did not match
Enter fullscreen mode Exit fullscreen mode

Running the Test

To run this test, navigate to the project directory in your terminal, in this case python_testing, and run:

pytest test_main.py
Enter fullscreen mode Exit fullscreen mode

Note it is pytest, not python. Running this shows pytest looking in the current directory, finding the test file, and reporting that one test case passed, along with how long it took.

Watching a Test Fail

It is worth seeing what a failure actually looks like, since you will be reading these messages often. If the assertion is changed to check for the wrong value, say, asserting the result equals "cold" instead of "hot", running the test again produces an AssertionError. Pytest reports exactly which test function failed, and shows that "hot" was not equal to "cold", pinpointing precisely what did not match.

That is the entire loop: write code, write an assertion about what it should do, run pytest, and get a clear pass or fail with details. You can also write multiple assert statements inside a single test function to check several things at once.


What a Unit Test Actually Is, and Why You Write Them

Now that you have written and run one test, it is worth stepping back and understanding what category of test this actually is, and why this practice matters.

A unit test is the smallest type of test. It targets one very small, isolated piece of code, typically a single function or method, sometimes a class. The purpose of a unit test is to confirm that this one small unit of code produces the expected result, in isolation from everything else in the codebase.

The value of this becomes clear once a project grows beyond a handful of functions. If something breaks in a large application, and you have good unit test coverage, you know almost immediately which small component is responsible, instead of hunting through a large, interconnected system to find the source of the problem. Unit tests isolate failures to a specific function or component that you can then go fix directly.

There are other categories of tests too, integration tests, system tests, end to end tests, each serving a different purpose, usually testing how multiple components work together rather than one piece in isolation. Unit tests specifically focus on the smallest building blocks.

Test Driven Development

Good test coverage is also useful during active development, since it is easy to accidentally break something you did not intend to touch. With solid coverage across a codebase, you can quickly see where something broke and fix it, especially if the tests themselves were written well.

This leads to an entire approach to software development called Test Driven Development, or TDD: writing the tests before writing the actual code. The idea is to define the requirements a function needs to satisfy first, in the form of test cases, and then write code specifically to make those tests pass, rather than writing code first and retrofitting tests to match whatever you happened to build.

If you have ever solved a coding problem on a platform like LeetCode, you have already experienced this pattern in a small way. The test cases are given to you upfront, and your job is to write a solution that satisfies all of them.


Writing More Assertions: Multiple Inputs and Edge Cases

Real functions usually need more than one test case to be properly covered. Consider two more functions in main.py:

# main.py

def add(a, b):
    return a + b

def divide(a, b):
    if b == 0:
        raise ValueError("Cannot divide by zero")
    return a / b
Enter fullscreen mode Exit fullscreen mode

Testing with Multiple Inputs

For add, a single test case is not enough to build confidence that the function is genuinely correct. You want to check it against a range of different inputs, including ones that might reveal edge case bugs.

# test_main.py

from main import add, divide
import pytest

def test_add():
    assert add(2, 3) == 5, "2 + 3 should equal 5"
    assert add(-1, 1) == 0, "negative one plus one should equal 0"
    assert add(100, 0) == 100, "100 + 0 should equal 100"
Enter fullscreen mode Exit fullscreen mode

A few things worth noting here. The function tests three different input pairs: (2, 3), (negative one, 1), and (100, 0), a normal case, a case involving a negative number, and a case involving zero. Each assert also includes an optional description string after the comma. This is not required, but it gives you a clearer message if that specific assertion fails, which is especially useful once you have several assertions inside one test function.

When you are testing a real function, the goal is to cover as many scenarios as reasonably possible, edge cases, empty inputs, unusual values, not just the obvious input you would expect someone to pass in. That is what makes a function's test coverage actually trustworthy rather than just a formality.

Testing That an Exception Gets Raised

The divide function is expected to raise a ValueError when dividing by zero, and that behavior deserves its own test, since whether a function fails correctly when it should is just as important as whether it succeeds correctly when it should.

def test_divide():
    assert divide(10, 2) == 5

    with pytest.raises(ValueError, match="Cannot divide by zero"):
        divide(10, 0)
Enter fullscreen mode Exit fullscreen mode

pytest.raises is a context manager that expects the code inside its with block to raise a specific exception, here ValueError. The match argument checks that the exception's message matches a given pattern (it works like a regular expression match against the error string). If divide(10, 0) does not raise a ValueError at all, or raises one with a different message, the test fails.

Running and Reading the Results

Running:

pytest test_main.py
Enter fullscreen mode Exit fullscreen mode

reports that the tests pass, along with the time taken.

If the match string is changed to something that does not correspond to the actual error message, say, checking for "can divide by zero" when the real message is "Cannot divide by zero", pytest fails the test and explicitly tells you that the regular expression pattern did not match the actual exception message, pointing to exactly where the mismatch is.

If multiple assertions fail at once, say, both an incorrect add result and a mismatched divide error pattern, pytest reports every individual failure separately, telling you specifically which assertion failed and why, so you are not left guessing which part of a larger test broke.


Fixtures: Setting Up Fresh State Before Every Test

As tests get more involved, especially ones involving objects with internal state, like classes, a new problem shows up: tests can accidentally affect each other if they are not properly isolated.

The Problem Fixtures Solve

Consider a small UserManager class:

# main.py

class UserManager:
    def __init__(self):
        self.users = []

    def add_user(self, name):
        if name in self.users:
            raise ValueError(f"User {name} already exists")
        self.users.append(name)
Enter fullscreen mode Exit fullscreen mode

You want to test two behaviors: that adding a user works correctly, and that adding a duplicate user correctly raises an error. Since these are two genuinely different behaviors, they belong in two separate test functions, each isolated to testing one specific thing. This keeps it easy to tell exactly what broke if either test fails, rather than bundling unrelated checks into one large test.

def test_add_user():
    user_manager = UserManager()
    user_manager.add_user("John Doe")
    assert "John Doe" in user_manager.users

def test_add_duplicate_user():
    user_manager = UserManager()
    user_manager.add_user("John Doe")
    with pytest.raises(ValueError):
        user_manager.add_user("John Doe")
Enter fullscreen mode Exit fullscreen mode

Notice that each test function creates its own fresh UserManager() instance. This matters more than it might look like at first. If both tests instead shared one global instance created outside the test functions, the first test's changes to that instance, adding John Doe, would carry over into the second test. The second test would then find a user manager that already contains John Doe for reasons unrelated to what it is actually trying to test, and it could fail for the wrong reason entirely, or pass by accident.

Tests need to run in isolation, against a consistent starting environment, not one that is silently shaped by whatever ran before it. Otherwise the order tests happen to run in could change whether they pass or fail, which defeats the purpose of testing in the first place.

Defining a Fixture

Rather than manually writing UserManager() at the top of every single test function, pytest provides fixtures, a way to define a setup step that runs automatically before each test that needs it.

# test_main.py

import pytest
from main import UserManager

@pytest.fixture
def user_manager():
    return UserManager()

def test_add_user(user_manager):
    user_manager.add_user("John Doe")
    assert "John Doe" in user_manager.users

def test_add_duplicate_user(user_manager):
    user_manager.add_user("John Doe")
    with pytest.raises(ValueError):
        user_manager.add_user("John Doe")
Enter fullscreen mode Exit fullscreen mode

The @pytest.fixture decorator marks the user_manager function as a fixture. Each test function that wants a fresh instance simply includes user_manager as a parameter, with the exact same name as the fixture function. Pytest sees that parameter name, recognizes it matches a defined fixture, runs the fixture function, and passes its return value into the test.

The result: every single test that requests this fixture gets its own brand new UserManager() instance, created fresh right before that specific test runs, never shared, never carried over from a previous test.

What Happens Without a Fixture

To make this concrete: if the fixture is removed and replaced with one global instance created once, outside of any test function,

user_manager = UserManager()  # created once, shared by everything
Enter fullscreen mode Exit fullscreen mode

running the test suite now produces a genuine failure. The first test (test_add_user) adds John Doe to the shared instance. By the time the second test (test_add_duplicate_user) runs, it is working with a UserManager that already has John Doe in it before the test's own logic even starts, because the first test never got cleared. The duplicate user check fails, not because the code is wrong, but because the tests were not properly isolated from each other.

This is exactly the problem fixtures are built to prevent. You could work around it manually by resetting state at the start of every test yourself, but a fixture handles it automatically and consistently.


Fixtures with Teardown: Setup and Cleanup

Fixtures are not limited to just creating a fresh object before a test runs, they can also clean up after a test finishes. This matters most when a test interacts with something that needs explicit cleanup: a database connection, a temporary file, a network resource.

Consider a simple in memory database simulation:

# main.py

class Database:
    def __init__(self):
        self.users = {}

    def add_user(self, name, age):
        if name in self.users:
            raise ValueError(f"User {name} already exists")
        self.users[name] = age

    def get_user(self, name):
        return self.users.get(name)

    def delete_user(self, name):
        if name in self.users:
            del self.users[name]

    def clear(self):
        self.users = {}
Enter fullscreen mode Exit fullscreen mode

Even though this example keeps everything in memory, the same class could just as easily be wrapping a real database, SQLite, Postgres, MongoDB, whatever. When you are testing something backed by a real database, you specifically want to make sure it is cleared or reset between test runs, or you risk hitting the exact same leftover state problem covered above.

Using yield for Setup and Teardown

# test_database.py

import pytest
from main import Database

@pytest.fixture
def db():
    database = Database()
    yield database
    database.clear()

def test_add_user(db):
    db.add_user("Alice", 30)
    assert db.get_user("Alice") == 30

def test_add_duplicate_user(db):
    db.add_user("Alice", 30)
    with pytest.raises(ValueError):
        db.add_user("Alice", 30)

def test_delete_user(db):
    db.add_user("Alice", 30)
    db.delete_user("Alice")
    assert db.get_user("Alice") is None
Enter fullscreen mode Exit fullscreen mode

The yield keyword is what makes this a setup and teardown fixture instead of just a setup fixture. Everything written before the yield line runs as the setup step, before the test executes. The yield database line hands that database instance to whichever test requested this fixture, the test then runs using that instance. Once the test finishes, whether it passed or failed, execution resumes right after the yield line, and anything written there runs as the teardown, or cleanup, step.

Fixture function

  code before yield   -> setup, runs BEFORE the test

  yield value          -> test runs, using this value

  code after yield     -> teardown, runs AFTER the test finishes
Enter fullscreen mode Exit fullscreen mode

In this example, database.clear() runs after every test that uses this fixture, resetting the database back to empty before the next test gets its turn, regardless of what the previous test did to it. If this were a real database connection instead of an in memory dictionary, the teardown step is exactly where you would close the connection, delete a temporary file, or run whatever cleanup the resource actually needs.

Running this test suite executes all three tests, adding a user, rejecting a duplicate, and deleting a user, and each one gets a properly reset database to work with, thanks to the teardown step running automatically between them.


Parameterized Testing: Avoiding Repetitive Test Code

Sometimes you want to run the exact same test logic against many different inputs and expected outputs. Writing a separate, nearly identical test function, or a long chain of separate assert lines, for each input gets repetitive fast, and it is easy to make small mistakes copying and pasting similar lines.

Suppose you have a function that checks whether a number is prime:

# main.py

def is_prime(n):
    if n < 2:
        return False
    for i in range(2, int(n ** 0.5) + 1):
        if n % i == 0:
            return False
    return True
Enter fullscreen mode Exit fullscreen mode

Instead of writing something like assert is_prime(7) == True, then assert is_prime(8) == False, then assert is_prime(9) == False, repeated for every number you want to check, pytest offers a decorator specifically built for this: pytest.mark.parametrize.

# test_main.py

import pytest
from main import is_prime

@pytest.mark.parametrize("num, expected", [
    (2, True),
    (7, True),
    (8, False),
    (9, False),
    (1, False),
    (0, False),
    (17, True),
    (18, False),
])
def test_is_prime(num, expected):
    assert is_prime(num) == expected
Enter fullscreen mode Exit fullscreen mode

The first argument to parametrize is a string naming the parameters, here "num, expected", which has to match the parameter names the test function actually accepts. The second argument is a list of tuples, where each tuple provides one set of values for those parameters. Pytest then runs the entire test function once per tuple, substituting in that tuple's values each time.

"num, expected"   matches   def test_is_prime(num, expected):

(2, True)    one full run of the test with num=2, expected=True
(7, True)    another full run with num=7, expected=True
(8, False)   another full run with num=8, expected=False
... and so on for every tuple in the list
Enter fullscreen mode Exit fullscreen mode

Running this reports each parameter combination as its own individual test result. If, say, 18 is deliberately marked as expected to be True instead of False in the parameter list, pytest reports one failure among the passing tests, and specifically identifies which parameter combination failed.

This pattern is especially valuable once you are testing functions with many meaningful edge cases. You get thorough coverage without duplicating the actual test logic, and adding a new case is as simple as adding one more tuple to the list. It also scales naturally to functions taking more than one input: you would simply add more names to the parameter string (say, "num1, num2, expected") and add matching values to each tuple.


Mocking: Testing Code Without Its Real Dependencies

This is one of the more important, and more commonly misunderstood, testing concepts. Real code frequently depends on things that either are not available, are not reliable, or simply should not be involved when you are testing something else, a network API, a database connection, an external service.

Why Mocking Matters

Imagine testing a piece of frontend code that depends on a backend API. Spinning up the actual backend, and all of its dependencies, just to test a small piece of frontend logic is unnecessary overhead, and worse, it ties your frontend test's pass or fail status to whether some unrelated backend service happens to be working at that exact moment.

If the backend has a genuine bug, or is temporarily down, that should not cause your frontend unit test to fail, because the frontend code itself might be completely correct. The failure exists somewhere else, in a different component, and a good unit test should stay focused on the one thing it is actually responsible for testing.

The solution is to mock the dependency, replace the real API call, database connection, or other external dependency with a fake version that returns controlled, predictable data. This way, the test verifies your code's logic in isolation, without being affected by whether some external system is actually working.

Mocking an API Call

Here is a function that fetches weather data from an external API:

# main.py

import requests

def get_weather(city):
    response = requests.get(f"https://api.weather.com/{city}")
    if response.status_code == 200:
        return response.json()
    else:
        raise Exception("Failed to fetch weather data")
Enter fullscreen mode Exit fullscreen mode

This function depends on an API you do not control. Maybe it requires an API key, maybe it is occasionally down, maybe it changes behavior. None of that should determine whether your test passes, because none of that is what the test is actually meant to check. What the test should check is whether this function correctly returns the JSON data when the status code is 200, and whether it correctly raises an exception otherwise.

Testing this without mocking would mean actually sending a real network request to a real API every time you run your tests, which is slow, unreliable, and dependent on a service you do not control. Mocking replaces that real request with a fake one you control completely.

# test_main.py

from main import get_weather

def test_get_weather(mocker):
    mock_get = mocker.patch("main.requests.get")
    mock_get.return_value.status_code = 200
    mock_get.return_value.json.return_value = {
        "temperature": 25,
        "condition": "sunny"
    }

    result = get_weather("London")

    assert result == {"temperature": 25, "condition": "sunny"}
    mock_get.assert_called_once_with("https://api.weather.com/London")
Enter fullscreen mode Exit fullscreen mode

Walking through this. The test function accepts a mocker parameter, this becomes available automatically once pytest mock is installed. mocker.patch("main.requests.get") replaces the real requests.get function, specifically the one referenced inside main.py, with a fake version for the duration of this test. The string path matters, it is patching requests.get as it is accessed from within main, not requests.get in some abstract global sense.

mock_get.return_value.status_code = 200 configures the fake response object so that calling the mocked get() returns something with status_code equal to 200. mock_get.return_value.json.return_value = {...} configures what calling .json() on that fake response returns, since .json is itself a method (a function), it has its own return_value to configure.

The test then calls the real get_weather function as normal. Internally, it calls what it thinks is requests.get, but it is actually calling the mock, so no real network request happens.

Finally, two things get checked: that the function's result matches what is expected given the mocked response, and separately, that the mock was actually called, and called with the specific argument expected. mock_get.assert_called_once_with(...) confirms the function reached out to the fake API exactly once, with the URL you would expect.

That last check is worth calling out specifically, because it is checking something different from the return value. It is not just whether the function produced the right output, it is whether the function actually called the dependency the way it was supposed to.

Checking How Many Times a Mock Was Called

Beyond checking that a mock was called exactly once with specific arguments, pytest mock supports a wider set of checks that come up often once you are testing real API integrations. You can check the exact number of times a mock was called:

def test_get_weather_called_correctly(mocker):
    mock_get = mocker.patch("main.requests.get")
    mock_get.return_value.status_code = 200
    mock_get.return_value.json.return_value = {"temperature": 25, "condition": "sunny"}

    get_weather("London")
    get_weather("Paris")

    assert mock_get.call_count == 2
Enter fullscreen mode Exit fullscreen mode

mock_get.call_count gives you the total number of times the mocked function was actually called during the test, which is useful when a piece of code is expected to call an API more than once, for example inside a loop, or as part of retry logic.

Checking That a Mock Was Never Called

Sometimes what you actually want to confirm is the opposite: that a dependency was never reached at all. This comes up a lot with routing style code, where one condition should call an API and another condition should not.

# main.py

def get_weather_if_needed(city, use_cache):
    if use_cache:
        return {"temperature": 20, "condition": "cached"}
    response = requests.get(f"https://api.weather.com/{city}")
    return response.json()
Enter fullscreen mode Exit fullscreen mode
def test_get_weather_uses_cache(mocker):
    mock_get = mocker.patch("main.requests.get")

    result = get_weather_if_needed("London", use_cache=True)

    assert result == {"temperature": 20, "condition": "cached"}
    mock_get.assert_not_called()
Enter fullscreen mode Exit fullscreen mode

mock_get.assert_not_called() confirms that the real API function was never triggered at all, which is exactly what you would want to verify here, since the whole point of the cache path is to avoid calling the API in the first place. If the caching logic had a bug and called the API anyway, this assertion would fail and tell you immediately, even though the returned data might otherwise look correct by coincidence.

Between assert_called_once_with, call_count, and assert_not_called, you have enough tools to verify not just what an API mock returned, but exactly how and whether your code interacted with it, which is usually the part that actually matters when you are testing code built around an external API.

Mocking a Database Connection

The same idea applies directly to database code. Here is a small function using sqlite3:

# main.py

import sqlite3

def save_user(name, age):
    conn = sqlite3.connect("users.db")
    cursor = conn.cursor()
    cursor.execute(
        "INSERT INTO users (name, age) VALUES (?, ?)",
        (name, age)
    )
    conn.commit()
    conn.close()
Enter fullscreen mode Exit fullscreen mode

Testing this normally would mean actually connecting to a real SQLite database file, which is not necessarily available, or desirable, in a testing environment. Instead, the connect function itself gets mocked:

# test_database.py

from main import save_user

def test_save_user(mocker):
    mock_connect = mocker.patch("main.sqlite3.connect")
    mock_cursor = mock_connect.return_value.cursor.return_value

    save_user("Alice", 30)

    mock_connect.assert_called_once_with("users.db")
    mock_cursor.execute.assert_called_once_with(
        "INSERT INTO users (name, age) VALUES (?, ?)",
        ("Alice", 30)
    )
Enter fullscreen mode Exit fullscreen mode

Here, mocker.patch("main.sqlite3.connect") fakes the connection itself. mock_connect.return_value.cursor.return_value reaches into the mock's chain to get a fake cursor object, matching how the real code calls conn.cursor() to get a real cursor. From there, the test checks two things: that connect was called with the expected database filename, and that cursor.execute was called with the expected SQL statement and parameters.

Running pytest test_database.py executes this without ever touching a real database file, no users.db gets created, no real SQL runs. The test purely verifies that save_user's logic is correct: that it connects to the right place and issues the right command, regardless of whether a real database happens to be available in whatever environment the test runs in.

Putting It All Together

Across these examples, a consistent set of ideas keeps showing up, and they are worth holding onto as the actual takeaways rather than the specific syntax.

Name test files and functions with the test_ prefix, that is how pytest discovers what to run. A test is just a function with an assertion in it, assert checks a condition, and pytest reports pass or fail with details. Keep each test isolated to one behavior, two different things being tested belong in two different test functions. Use fixtures to guarantee a clean, consistent starting state for every test, with yield, a fixture can also clean up afterward. Use pytest.mark.parametrize when the same logic needs to run against many different inputs, instead of duplicating test code. Mock external dependencies, APIs, databases, anything outside your control, so a unit test only ever fails because of the code it is actually meant to be testing, not because of something unrelated going wrong elsewhere. When mocking an API, check both what your code returned and whether it actually called the dependency the way it was supposed to, using assert_called_once_with, call_count, or assert_not_called depending on what you need to verify.

None of this requires memorizing pytest's entire feature set. Most real world test suites are built almost entirely out of these ideas, combined and reused across a project. Once fixtures, parametrize, and mocking feel natural, you have what you need to write solid test coverage for the kinds of Python projects you will actually build.

Top comments (0)