Let's say you want to know how many times the word research appears in a document.
You could upload the document to a program and read the file one byte at a time, compare each byte with the first letter of the word, and inspect the promising positions more closely.
Now give the same program ten megabytes of text. Then a hundred megabytes. Then give it the papers, policy reports, interview transcripts, and field notes collected by an entire research team.
The question has not changed. The computer is still asking, “Does the word start here?” The problem is how many times it has to ask.
This is where SIMD starts to make sense.
SIMD means single instruction, multiple data. Instead of comparing one byte and then the next one, the processor can compare several bytes together. The values sit in separate lanes of a vector, and the same operation runs across those lanes.
So SIMD is not really a trick for making a loop move faster. It changes the shape of the work. A scalar loop moves through the data one value at a time. A SIMD operation looks across several values at once.
That sounds simple, and the basic idea is simple. The interesting part begins when we try to use it in a real program.
What happens when the file does not divide neatly into complete vectors? What happens when the first byte matches but the rest of the word does not? How do we distinguish research from researcher?
And after making the code faster, how do we prove that it still returns the right answer?
We will answer those questions by building ResearchLens, a Mojo application that searches research text. It uses SIMD to find possible matches and a scalar verifier to decide which ones are real.
Along the way, we will look at:
- what a SIMD lane is
- why the vector width is part of a Mojo type
- how loads, comparisons, masks, and reductions fit together
- why a fast candidate is not automatically a correct match
- how to handle the final partial vector safely
- how to benchmark the result without making claims the data cannot support.
You do not need a GPU for this tutorial. The SIMD hardware is already inside the CPU in a modern computer. We are going to show Mojo how to use it.
The scalar loop
Let us begin with the obvious version.
def count_first_byte_scalar(data: Span[UInt8, _], first: UInt8) -> Int:
var count = 0
for index in range(len(data)):
if data[index] == first:
count += 1
return count
This function is easy to trust. It checks one byte, updates the count, and moves on.
If the input contains eight bytes, the loop asks the same small question eight times:
Is this byte equal to the first byte of the term?
The important word is same. Every comparison is independent. The answer for byte 3 does not depend on the answer for byte 2.
That is the kind of work SIMD is built for.
SIMD does not remove the work. It presents the work to the processor in a shape the hardware can execute more efficiently.
A register with lanes
A SIMD value is one fixed-width value divided into lanes. Every lane contains the same type of element, and an operation is applied element by element.
For example, this value contains eight 16-bit integers:
var values = SIMD[DType.int16, 8](1, 2, 3, 4, 5, 6, 7, 8)
var squared = values * values
print(squared)
# [1, 4, 9, 16, 25, 36, 49, 64]
The multiplication looks ordinary, but its inputs are vector-shaped. Each lane is multiplied by the corresponding lane.
The type tells us two things:
SIMD[DType.int16, 8]
-
DType.int16tells Mojo what each lane contains. -
8tells Mojo how many lanes the vector has.
So this is one SIMD value containing eight 16-bit integers. Both values are compile-time parameters. By the time Mojo generates the machine code, it already knows the element type and the number of lanes.
Mojo takes this model further than you might expect. Types such as Float32, Int64, and UInt8 are scalar aliases built on the same SIMD foundation. A scalar is a one-lane SIMD value.
So Mojo does not give us one disconnected model for scalar numbers and another for vectors. It gives us one model that grows from one lane to many.
Note: One vector-shaped operation does not mean one clock cycle. The compiler may map a value to one hardware register, several registers, or another efficient representation. A wider type is not automatically faster.
The useful width depends on the element size, the target CPU, memory behaviour, and the rest of the kernel. Vector width is something to measure, not assume.
The SIMD data path
Most useful SIMD kernels are not one clever instruction. They are a short pipeline.
The stages are separate for a reason:
| Stage | Input | Vector work | Correctness rule |
|---|---|---|---|
| Load | Contiguous bytes | Read a complete vector | Never read beyond the buffer |
| Compare | First target byte | Produce one Boolean per lane | Use an element-wise comparison |
| Select | Boolean mask | Map true and false to 1 and 0
|
Keep the mask inspectable |
| Reduce / Iterate | Candidate lanes | Combine lane values | A candidate is not yet a complete match |
| Verify | Candidate positions | Check the full term | Match the scalar reference |
This is the architecture we will use for the ResearchLens project.
Why ResearchLens uses two algorithms
ResearchLens has one job: count exact ASCII terms in UTF-8 research text.
It uses two implementations:
- A scalar implementation that is intentionally obvious.
- A SIMD implementation that scans several possible starting positions at once.
The scalar version is not code that we keep around for nostalgia. It is the correctness oracle.
If the SIMD version and the scalar version disagree, the SIMD version is wrong.
The ResearchLens project has three areas:
Data layer
Research papers, reports, transcripts, and field notes become a contiguous byte buffer.
The first version is deliberately byte-oriented. UTF-8 text can contain multi-byte characters, but ASCII letters have the same byte representation inside UTF-8. That makes exact ASCII terms a useful, bounded starting point.
It does not make ASCII boundary rules a complete Unicode tokenizer. We will say that plainly instead of hiding it in a footnote.
Decision layer
The SIMD scan checks the first byte of the target term across a complete vector.
If the term is research, each byte equal to r becomes a candidate position.
That narrows the search. It does not prove that the complete word exists.
Verification layer
The verification layer includes a scalar function that checks every candidate:
- Does the complete term fit inside the buffer?
- Do all bytes match?
- Is the byte before the term a word character?
- Is the byte after the term a word character?
This is where the program distinguishes research from researcher, or art from partial.
SIMD narrows the search while the verifier owns the meaning.
Building the verifier
Before we write the vector loop, we need the function that defines a correct match.
def is_ascii_word_byte(byte: UInt8) -> Bool:
return (
(byte >= 48 and byte <= 57)
or (byte >= 65 and byte <= 90)
or (byte >= 97 and byte <= 122)
or byte == 95
)
def term_matches_at(
data: Span[UInt8, _], term: Span[UInt8, _], start: Int
) -> Bool:
if len(term) == 0 or start < 0 or start + len(term) > len(data):
return False
for index in range(len(term)):
if data[start + index] != term[index]:
return False
if start > 0 and is_ascii_word_byte(data[start - 1]):
return False
var end = start + len(term)
if end < len(data) and is_ascii_word_byte(data[end]):
return False
return True
This function is not vectorized, and that is fine because the SIMD scan should make the set of positions reaching this function much smaller than the full input. The verifier can then spend more work on the positions that deserve it.
It's a common performance flow to use a cheap parallel filter, then apply a more expensive exact check.
Keep the scalar reference obvious
Now we can write the reference implementation.
def count_term_scalar(data: Span[UInt8, _], term: Span[UInt8, _]) -> Int:
if len(term) == 0 or len(term) > len(data):
return 0
var count = 0
var limit = len(data) - len(term) + 1
for start in range(limit):
if term_matches_at(data, term, start):
count += 1
return count
The SIMD implementation has more moving parts, so we need a simple version to test it.
Load, compare, and inspect the mask
Here is the centre of the SIMD path in the src/analyzer.mojo file:
comptime SIMD_WIDTH = 32
var first = term[0]
var ptr = data.unsafe_ptr()
var chunk = ptr.unsafe_load[width=SIMD_WIDTH](offset)
var candidates = chunk.eq(
SIMD[DType.uint8, SIMD_WIDTH](first)
)
The first line in the code to look at is:
var chunk = ptr.unsafe_load[width=SIMD_WIDTH](offset)
It loads 32 adjacent bytes from data, beginning at offset, into one SIMD[DType.uint8, 32] value.
The next line compares all 32 bytes with the first byte of the search term:
var candidates = chunk.eq(
SIMD[DType.uint8, SIMD_WIDTH](first)
)
The result is a SIMD mask. Each lane tells us whether the corresponding byte could be the beginning of a complete match.
The name is deliberately honest. Mojo cannot prove that a raw pointer has 32 valid bytes available at this offset.
The loop condition must establish that contract before the load happens.
Then .eq() compares every lane with the first byte of the target.
The explicit method matters. In Mojo, == answers whether two complete SIMD values are equal and returns one Bool. The element-wise .eq() method returns a Boolean SIMD mask, which is what we need here.
Conceptually, the result looks like this:
bytes: m o j o r e s e a r c h
first byte: r r r r r r r r r r r r r
mask: F F F F F T F F F F T F F
The true lanes are candidate starting positions. We still have to verify the full term at those offsets.
The complete SIMD loop
The complete function adds the verifier and the tail:
comptime SIMD_WIDTH = 32
def count_term_simd(data: Span[UInt8, _], term: Span[UInt8, _]) -> Int:
if len(term) == 0 or len(term) > len(data):
return 0
var count = 0
var limit = len(data) - len(term) + 1
var offset = 0
var first = term[0]
var ptr = data.unsafe_ptr()
while offset + SIMD_WIDTH <= limit:
var chunk = ptr.unsafe_load[width=SIMD_WIDTH](offset)
var candidates = chunk.eq(
SIMD[DType.uint8, SIMD_WIDTH](first)
)
for lane in range(SIMD_WIDTH):
if candidates[lane] and term_matches_at(
data, term, offset + lane
):
count += 1
offset += SIMD_WIDTH
while offset < limit:
if data[offset] == first and term_matches_at(data, term, offset):
count += 1
offset += 1
return count
The loop does not reduce the mask to one number because we need the actual candidate positions. It reads each Boolean lane and sends only the true positions to the verifier.
Once the tests are stable and the benchmark tells us verification is the bottleneck, we have a precise place to improve.
The boundary is part of the algorithm
A document length is rarely divisible by 32.
Suppose 103 bytes are valid starting positions. The vector loop can safely process three complete 32-byte windows. That covers 96 positions. Seven remain.
An unchecked load at offset 96 would try to read 32 bytes when only seven valid positions remain. That is not a performance issue. It is an invalid memory access.
The loop guard makes the contract explicit:
while offset + SIMD_WIDTH <= limit:
# A complete vector is safe here.
...
while offset < limit:
# The remaining positions use the scalar path.
...
The code above is because limit is not greater than len(data), the first condition also ensures that the 32-byte unsafe_load stays within the input buffer.
The scalar tail is part of the algorithm. A vector can tell us that a lane contains r. It cannot, by itself, tell us that research is a complete word in the application’s tokenization model.
The fast path and the correctness path need a clean boundary between them.
Overview of the data shapes
When you read a SIMD kernel, asking “which line is fast?” is usually less helpful than asking “what shape does the data have here?”
The ResearchLens path is:
- Load: 32 adjacent bytes become one typed vector.
- Compare: 32 bytes become 32 Boolean answers.
- Mask: the answers become inspectable candidate lanes.
- Reduce or iterate: the vector result becomes a smaller scalar decision.
- Verify: a candidate position becomes a correct application-level match or a rejection.
That is the mental model to keep. Each line changes the shape of the data.
Setting up the project in Mojo
We have discussed how the SIMD path works. Now let us set up the complete project and run it.
ResearchLens was built and tested with Mojo 1.0.0. You can write Mojo in VS Code, Windsurf, or another editor that supports the Mojo extension.
Install Mojo with Pixi
The official Mojo installation guide supports both uv and Pixi for creating and managing a Mojo project. We will use Pixi because it creates a reproducible project environment and generates a lockfile for the dependencies.
If Pixi is not already installed, run:
curl -fsSL https://pixi.sh/install.sh | sh
Restart your terminal after the installation if the pixi command is not immediately available.
Now create the project:
pixi init researchlens \
-c https://conda.modular.com/max/ \
-c conda-forge
cd researchlens
pixi add "mojo==1.0.0"
pixi run mojo --version
The final command should print the installed Mojo version:
Mojo 1.0.0 (ed45d567)
The value in parentheses,
ed45d567, is the build identifier for this specific Mojo 1.0.0 build.
Mojo may have a newer release by the time you read this. We are using version 1.0.0 because that is the exact version used to build, test, and benchmark this project.
Create the project files
ResearchLens has a small structure:
researchlens/
├── pixi.toml
├── pixi.lock
├── README.md
├── data/
│ └── sample.txt
├── src/
│ └── analyzer.mojo
└── tests/
└── test_analyzer.mojo
Create the directories:
mkdir -p data src tests
Add the source code to src/analyzer.mojo, the tests to tests/test_analyzer.mojo, and a small research-text sample to data/sample.txt.
The complete project is available in the ResearchLens repository.
Configure the Pixi tasks
Open pixi.toml and use the following configuration:
[workspace]
authors = ["Your Name"]
channels = ["https://conda.modular.com/max", "conda-forge"]
name = "researchlens"
platforms = ["linux-64", "linux-aarch64", "osx-arm64"]
version = "1.0.0"
[dependencies]
mojo = "==1.0.0"
[tasks]
analyze = "mojo run src/analyzer.mojo"
test = "mojo run -I src tests/test_analyzer.mojo"
The two tasks give us shorter commands:
pixi run test
pixi run analyze data/sample.txt research 100
Pinning the Mojo version matters. It means another engineer can clone the repository, install the same dependencies, and reproduce the environment we tested.
Commit pixi.lock to the repository as well. The configuration describes the environment, while the lockfile records the exact packages Pixi resolved.
Test the scalar and SIMD paths together
A fast implementation is only useful if it returns the correct answer.
Instead of testing the scalar and SIMD functions separately, we give both functions the same input and confirm that they return the same match count.
Open tests/test_analyzer.mojo:
from std.testing import assert_equal
from analyzer import count_term_scalar, count_term_simd
def check(text: String, term: String, expected: Int) raises:
var data = text.as_bytes()
var needle = term.as_bytes()
assert_equal(count_term_scalar(data, needle), expected)
assert_equal(count_term_simd(data, needle), expected)
def main() raises:
check("", "research", 0)
check("research", "research", 1)
check("research research", "research", 2)
check("researcher research", "research", 1)
check(
"A long prefix that crosses a vector tail: research",
"research",
1,
)
print("ResearchLens: all scalar/SIMD agreement tests passed")
Run the tests:
pixi run test
The result should be:
ResearchLens: all scalar/SIMD agreement tests passed
Notice that check is marked with raises:
def check(text: String, term: String, expected: Int) raises:
Some of the operations called inside check may return an error. Mojo requires the function to make that possibility explicit.
Without raises, the compiler reports that a function which may raise is being called from a context that cannot raise. We encountered this error while building ResearchLens, and adding raises fixed it.
Run the analyzer
With the tests passing, run ResearchLens against the sample corpus:
pixi run analyze data/sample.txt research 100
The three arguments are:
| Argument | Meaning |
|---|---|
data/sample.txt |
The text corpus to search |
research |
The term to find |
100 |
The number of benchmark iterations |
On the machine used for this tutorial, the command produced:
bytes: 244
matches: 3
iterations: 100
scalar total ns: 64000
SIMD total ns: 19000
ResearchLens found three complete occurrences of research. It then ran both implementations 100 times and measured their total execution time.
For this run, the SIMD implementation was approximately 3.37 times faster than the scalar implementation.
Build the executable
So far, Pixi has been running the Mojo source file directly. We can also compile ResearchLens into a standalone executable:
pixi run mojo build src/analyzer.mojo -o researchlens
Run the compiled application:
./researchlens data/sample.txt research 100
On the Apple Silicon machine used for this tutorial, Mojo produced an 86 KB Mach-O arm64 file.
You do not need a GPU for this project. We are using CPU SIMD. The vector hardware is already inside a modern processor, and Mojo gives us a way to express work that can run across those vector units.
Measure the complete application
A vector width of 32 does not exactly mean a 32-times speedup.
The program still has to:
- Load the bytes from memory.
- Find candidate positions.
- Verify each candidate.
- Handle the final partial vector.
- Count the valid matches.
The compiler may also divide a wide SIMD value into smaller hardware operations. The final performance depends on the processor, the data, the compiler, and the way the algorithm is written.
That is why we benchmark the complete application instead of timing only the most impressive SIMD instruction.
For a useful comparison:
- Run the scalar and SIMD implementations on the same bytes.
- Confirm that they return the same match count.
- Keep compilation and one-time setup outside the timed section.
- Run enough iterations to reduce timing noise.
- Record the CPU, operating system, Mojo version, corpus size, search term, vector width, and iteration count.
Here are the results from the completed ResearchLens project:
| Corpus | Iterations | Matches | Scalar total | SIMD total | Speedup |
|---|---|---|---|---|---|
| 244 bytes | 100 | 3 | 64,000 ns | 19,000 ns | 3.37× |
| 10,485,996 bytes | 100 | 127,878 | 2,745,454,000 ns | 856,933,000 ns | 3.20× |
| 104,857,746 bytes | 20 | 1,278,753 | 5,531,949,000 ns | 1,732,811,000 ns | 3.19× |
The 10 MB and 100 MB tests settled at almost the same ratio: approximately 3.2 times faster for this workload on this machine.
That consistency is more useful than the result from the 244-byte sample. With such a small file, timer noise and fixed overhead account for a larger part of the measurement.
The 100 MB result does not mean Mojo will make every program 3.19 times faster. It means this implementation measured 3.19 times faster with this:
- Search term
- Corpus
- Vector width
- Mojo version
- Processor
- Number of iterations
Change any of those variables and the result may change.
The data also matters
ResearchLens first uses SIMD to find bytes that could begin a match. It then sends those candidate positions to the scalar verifier.
We can describe that relationship with:
candidate ratio =
candidate starting bytes / valid starting positions
If the search term begins with an uncommon byte, the SIMD filter can reject most positions cheaply.
If it begins with a common byte, more positions reach the scalar verifier, and verification takes a larger share of the execution time.
The value of the optimization therefore depends partly on the data being searched.
Next steps
SIMD is useful when the same small operation must be applied to many independent values.
Mojo makes the parts of that work visible in ordinary code:
- The element type
- The lane count
- The vector load
- The comparison mask
- The reduction
- The scalar verifier
- The tail
We can inspect each part, test it, and measure it.
ResearchLens is a small application, but the same pattern appears in delimiter scanning, byte classification, image processing, numerical transformations, filters, checksums, and many other workloads.
The goal is not to vectorize everything.
It is to notice when the computer is being asked the same small question millions of times.
If you build your own version of ResearchLens, try changing the vector width or the distribution of the first byte and share in the comments.
P.S: If you want to learn specifically about SIMD, read Mitchell Hashimoto's article.





Top comments (0)