DEV Community

Sergei Karmanovich
Sergei Karmanovich

Posted on Originally published at habr.com AI-assisted

One Mock, Three Problems

In this article I describe three problems I ran into while using mocks:

  • Lost invariants
  • Requirements that live only in your head
  • Mute tests

A bit of background first.

Why I disliked mocks

For a long time I disliked mocks. Specifically Mock — not Fake, Stub, Spy, or the other kinds of test doubles.

Mostly because I kept seeing them used as a substitute for almost everything:

I am testing class X. I do not care that it uses class Y: other tests are supposed to check Y.

There is some truth in that. A test of class X does not have to re-check every branch of class Y.

It does not follow that the rules of class Y can be ignored completely.

At the time I did not know enough to explain, in concrete terms, why mocking everything felt wrong. So I went deeper into testing and read Vladimir Khorikov's Unit Testing Principles, Practices, and Patterns.

That book helped me see the difference between the classic and London schools of unit testing.

Simplified:

  • the classic school isolates tests from each other;
  • the London school tries to isolate the object under test from its dependencies.

After the first reading, the London approach felt natural. In our services the main dependencies were repositories and clients of external systems. We replaced repositories with Fake implementations and clients with mocks.

For those boundaries, it worked.

The problems showed up later, with more complex scenarios: application services, domain services, calculators, managers, and other classes that contained business rules.

That is when I could finally say why mocking your own application classes is dangerous.

The main reason is lost invariants.

1. Lost invariants

Consider a simplified, fictional example.

We have a service that creates a user. It enforces one business rule: the user must be at least 18.

class UserServiceApp:
    def __init__(self, user_repository: BaseUserRepository) -> None:
        self._user_repository = user_repository

    def create(self, user_data: CreateUserDTO) -> None:
        self._check_user_data(user_data)
        self._user_repository.add(user_data)

    def update(self, user_data: UpdateUserDTO) -> None:
        self._check_user_data(user_data)
        self._user_repository.update(user_data)

    def _check_user_data(self, user_data: CreateUserDTO | UpdateUserDTO) -> None:
        if user_data.age < 18:
            raise DomainLogicError("We only work with adults.")
Enter fullscreen mode Exit fullscreen mode

This class is easy to test if you pass it a Fake repository.

Yes, the age check can — and should — move into a factory method on the entity itself. I leave it in the service on purpose, so the problem stays visible.

We will come back to that objection.

So far everything looks fine:

  • the system has a rule: a user must be 18;
  • the rule is enforced;
  • tests cover it.

How many other scenarios use this service?

In a large system, creating or updating a user can be called from dozens of use cases:

class RegisterUserUseCase: ...        # the user registers themselves
class AdminCreateUserUseCase: ...     # a manager creates the user in the admin panel;
                                      # the logic around creation is different
class AcceptInvitationUseCase: ...    # a company invitation:
                                      # the user is created when they accept
class CreateReferralUserUseCase: ...  # registration via a referral link
...
# And that is only creation. Profile updates, email changes,
# blocking, imports — separate scenarios.
# In every test, a UserServiceApp mock throws the same rules away again.
Enter fullscreen mode Exit fullscreen mode

In the test we usually create a mock of the service and check that it was called with the expected arguments. What if the test data is a 17-year-old user:

def test_create_user() -> None:
    user_service_app_mock = Mock(spec=UserServiceApp)
    use_case = CreateUserUseCase(user_service_app=user_service_app_mock)

    use_case.execute(user_data=CreateUserData(age=17, name="Ivan"))

    user_service_app_mock.create.assert_called_once_with(
        CreateUserData(age=17, name="Ivan"),
    )
Enter fullscreen mode Exit fullscreen mode

We run the tests, and they pass.

============================= test session starts ==============================
platform linux -- Python 3.x.x, pytest-x.y.z, pluggy-x.y.z
rootdir: /path/to/project
collected 1 item

test_....py .                                                            [100%]

============================== 1 passed in 0.02s ===============================
Enter fullscreen mode Exit fullscreen mode

Formally this is correct: the use case called its dependency with the expected arguments, and the mock confirmed the call.

From the system's point of view, the test is lying.

In production that call fails: the real UserServiceApp does not work with users under 18. The mock knows nothing about the rule and accepts any data.

We checked that the classes interacted. We lost whether that interaction is allowed.

This is not limited to a user's age. A real system has dozens or hundreds of rules:

  • an order cannot be placed when there is an outstanding debt;
  • a user's email must be unique;
  • an operation amount must not exceed the available limit;
  • an entity's status must allow the transition;
  • the same resource cannot be reserved twice;
  • an operation must take another aggregate's state into account.

Some of these rules can and should live inside entities. Not every rule fits in one object's constructor or factory. Some invariants depend on other entities, on storage, or on a domain service.

When we replace such an object with a mock, its rules disappear from the test.

Every mock of your own class that holds business rules can steal invariants from the scenario under test.

The test stays green, but it no longer guarantees that the call it checks is even valid in the real system.

That is the first problem: lost invariants.

2. Requirements that live only in your head

Suppose we understand the problem and still do not want to change the approach. Instead of using the real service, we decide to watch, by hand, that only valid data reaches the mock.

We walk through the tests and change user ages to 18 or above.

We might even add a test factory:

make_adult_user()

Now the minimum age is no longer scattered across dozens of tests.

If that were the only invariant in the system, this might be enough.

A real application has many more rules. Are we going to build a separate factory for each of them?

Even if we do, the developer still has to know:

  • which factory to pick;
  • which rules the called dependency enforces;
  • which combinations of data are allowed;
  • which rules the factory already covers, and which it does not.

We can put the most experienced developers in charge. Let them review every pull request and check that every mock receives valid data.

How long will that last?

We split the system into classes and distributed responsibility across objects so that nobody would have to hold the whole system in their head at once.

After we mock our own classes, that responsibility does not go away. It moves out of executable code and into the developer's memory.

The class no longer protects the test with its own checks. A person has to remember those checks instead.

We still use OOP, encapsulation, and separation of responsibility — and in tests we voluntarily give up a large part of what they are for.

The larger the system, the more context you have to hold.

The test stops using the system's knowledge and starts depending on the test author's knowledge.

That is the second problem: requirements that live only in your head.

3. Mute tests

Now imagine the team manages to keep the test data correct by hand for a while.

There is another problem, and it is not obvious to everyone: business rules change.

Today the minimum age is 18. Tomorrow, because of a law or a new requirement, it becomes 21.

We change the check in UserServiceApp.

The service's own tests will most likely fail. They use the real object, so they see the change in its behavior.

What happens to the use-case tests where that service is replaced by a mock?

Nothing.

They keep passing users who are 18, and they stay green.

The mock does not know that the real class's contract changed. There is no feedback between the mock and the real object.

In the best case, the team remembers the shared factory and updates it. That only works if someone knows the factory exists and which tests depend on it.

In the worst case, the data is hardcoded in dozens of tests. Part of the suite keeps describing behavior the system no longer has.

These tests are dangerous not because they fail, but because they do not fail when they should.

They look current. They pass in CI. A developer opens them to understand how the system behaves and gets outdated information.

The tests cannot see the change in the real class, so they stay silent.

That is the third problem: mute tests.

Does this mean you should never use mocks?

No.

The problem is not Mock itself. It is the boundary where you use it.

A mock is a good fit for dependencies that live outside the process or the system:

  • an HTTP client of a third-party API;
  • an email provider;
  • a payment gateway;
  • a message broker;
  • and so on.

You do not need to run real external systems in every unit test.

If the dependency is your own domain service, or an application class that contains business rules, it is often safer to leave it real.

The infrastructure underneath can still be replaced with controlled test implementations:

  • repository — Fake;
  • external client (or, better, its transport) — Mock;
  • domain service — real;
  • scenario under test — real.

Then the test uses the same business rules as production code, and it stays fast and deterministic.

Conclusion

One mock of your own class that contains business logic can create three problems at once:

  1. Lost invariants — the mock accepts data the real object would reject.
  2. Requirements that live only in your head — the developer has to remember the mocked class's rules and reproduce them in tests by hand.
  3. Mute tests — when the real contract changes, the dependent tests keep passing.

So before I create another mock, I ask myself:

Am I replacing an external boundary of the system, or my own object that knows something important about the business?

If it is the second one, the mock does not simplify the test.

It just makes the problem invisible.


Originally published in Russian on Habr.

Top comments (0)