DEV Community

Dennis
Dennis

Posted on

Improve code quality in dotnet using AI and code analysis

The amount of code that is written has significantly increased with the rise of LLMs. Along with it, grows uncertainty about the quality of our work. How do we know that the code written by AI is good? How can we trust that a non-deterministic process generates high-quality code without bugs? In the practice of Harness-engineering, this uncertainty is addressed with sensors: deterministic tools that validate the output of LLMs (or people) using consistent rules.

In this blog I propose one possible answer to the question: how do we build such sensors in dotnet applications? I will tell about my experience with building code analyzers and give some general tips for when you want to dive into the topic for yourself.

Roslyn analyzers

One option is to write roslyn analyzers. Roslyn is the name for the C# compiler that the dotnet CLI and visual studio use to build your source code. If you've enabled nullable reference types or you've installed StyleCop before, then you're already familiar with what roslyn analyzers do: They identify problematic code patterns and give you warnings and errors at build-time.

As it turns out, LLMs are very good at writing roslyn analyzers. It's also very easy to scaffold a code analyzer project using Microsoft's instructions. I got my first simple working analyzer up in about 3 hours. Custom roslyn analyzers have these benefits and drawbacks:

  • ✅ Easy to make with LLMs: With a good description of what you want, LLMs are excellent at writing custom analyzers.
  • ✅ Quick to get started with: Setting up a project with code analyzers is easy to do.
  • ✅ Low risk: The code for analyzers does not run in production, only on your machine during development. Code quality for analyzers is therefore not as important as otherwise. If it doesn't do what you need, you can throw it away without losing anything.
  • ✅ Also benefits development without LLMs: Recognizing bad code patterns is good no matter how you develop. Being able to enforce good practices is a super power no matter what.
  • ✅ Exposes blindspots: Code analyzers are deterministic and will therefore always work the same. If you failed to enforce your rule anywhere in your code, then an analyzer will find it, guaranteed.
  • ❌ Cannot co-exist with source code: Code analyzers are either installed as NuGet package or as Visual Studio plugin. You cannot add analyzers as a class library alongside your source code. You can therefore also not evolve it together with your source code. You must separately publish a NuGet package or Visual Studio plugin.
  • ❌ May significantly increase warnings in code: If a code analyzer rule is invasive, it may introduce many new warnings. If you have warnings-as-errors enabled, then a code analyzer that introduces many new warnings may not be productive. A long list of warnings may hide other important issues.
  • ❌ All warnings must be fixed: You can only enforce the rules that you make by failing the build. If you don't do this, the LLM may choose to ignore your warnings.

When are Roslyn analyzers a good choice?

Roslyn analyzers work on the syntax tree of your code. While they are good at identifying patterns in syntax, they are not great at inferring intent. In other words: It's good at seeing what you do, but not good at understanding why you do it.

An example of a good use-case is as follows:
A reference to a disposable object should not escape the scope of the using block in which the object is created.

using (var myObject = new SomeDisposableType())
{
    return new ResultType(myObject); // WARNING: disposable object should not escape the scope of it's corresponding using-block.
}
Enter fullscreen mode Exit fullscreen mode

This is a concrete syntax pattern with demonstrable negative effects: consumer code may throw ObjectDisposedException. A Roslyn analyzer can recognize this syntax pattern and generate warnings for it.

On the other hand, a Roslyn analyzer is not good at inferring intent. For example:
Two or more mutually exclusive properties on a model should be replaced with a single discriminated union.

public class MyPageModel
{
    // WARNING: two mutually exclusive properties should be replaced with a single discriminated union
    public HeaderSmallVariant? HeaderSmall { get; init; }
    public HeaderLargeVariant? HeaderLarge { get; init; }
}
Enter fullscreen mode Exit fullscreen mode

This case is ambiguous: is it possible to have both a small and a large header at the same time? Is it possible to have neither? We can only infer this by looking at how the type is used. We need to use context clues to know what the writer of the code meant.

Tips to get started

Here are some broad steps and findings that will help you get up-and-running with analyzers.

  1. Follow the Microsoft instructions for setting up an analyzer project. The scaffolded project works out-of-the-box and comes with an example.
  2. Do not immediately switch the test-framework. You should update all packages first to the latest versions. You'll find that several packages need to be replaced. NuGet recommends alternatives and those alternatives are test-framework agnostic. After replacing the packages, you'll need to fix some build errors. The fixes are straightforward.
  3. Ask the LLM to replace MSTest with your test-framework of choice. The LLM can do this very well.
  4. Do not upgrade Microsoft.CodeAnalysis.CSharp and Microsoft.CodeAnalysis.CSharp.Workspaces to the latest version. If you're still developing applications on dotnet 8, version 4.3.0 works well. If you're on dotnet 10, you could also use 5.0.0, but 4.3.0 works for both. Upgrading to the latest version will cause additional warnings when you install your analyzer package in a project.
  5. If you have an example of bad code in your source repository, ask the LLM to export the sample into a document. Also find examples of better code. Make the LLM write an explanation of why the bad code is bad and the good code is good. Ask the LLM to fully qualify types in the document. When you think the document is complete, move over to your analyzer solution and feed it to the LLM. Ask it to plan an analyzer to find warnings for the bad code. This gets you the best results.
  6. Ask the LLM for examples in the form of unit tests. The unit tests for code analyzers are easy to read and illustrate how the analyzer works.

At the moment I still have one problem: I am unable to debug the analyzer and manually test it with the visual studio plugin. I don't know why. It hasn't been a big problem for me, because the unit tests already helped me solve all the problems that I had thus far.

Final thoughts

It was surprisingly easy to get started with code analyzers using LLMs. This experiment has made me think deeper about what code quality means and if it's possible to make it more measurable somehow. Although I think code analyzers have somewhat limited applicability, they do seem to be very helpful in some cases. Now that the bar has been significantly lowered to create code analyzers, I believe they are a viable option for everyday development tasks, with or without LLM support.

Thank you for reading and I'll see you in my next blog! 😊

Top comments (0)