Every Java team that touches local AI knows the same story. You build a clean Spring Boot or Quarkus service, then someone says "we need inference on our own GPU," and suddenly your tidy JVM stack has a Python sidecar bolted to the side. PyTorch, a CUDA toolkit, a second Dockerfile, a second set of CVEs, and a REST hop between the two halves of your own application.
On October 8, a project started making the rounds on r/LocalLLaMA with a claim that attacks exactly that pain: a Java framework that compiles Java bytecode down to CUDA and cuTile, running LLM inference on NVIDIA GPUs at a claimed 90% of llama.cpp performance. No C++ code, no Python process, no sidecar.
The project is jitLLM, built by the TornadoVM team at the University of Manchester with Red Hat as a collaboration partner. I have not run it on my own hardware yet, so this is a full-disclosure analysis piece, not a benchmark report. Everything below comes from the project's own pages, the TornadoVM blog, and the discussion around the announcement. But the engineering behind it is real, documented, and worth understanding even if you never install it.
What jitLLM actually is
A Java-native inference engine, not a wrapper. jitLLM (Java Inference Tornado toolkit) runs Llama 3, Mistral, Qwen 2.5, Qwen 3, Phi-3, IBM Granite, and Devstral 2 models in GGUF format. It does not shell out to llama.cpp or call Ollama over HTTP. The transformer inference loop itself is written in Java, and the heavy math gets JIT-compiled to GPU kernels at runtime.
The GPU compiler underneath is TornadoVM. This is the part that makes the claim even plausible. TornadoVM is an open-source plugin to the JDK that takes annotated Java bytecode and compiles it to OpenCL, PTX/CUDA, or SPIR-V. It has been under active development for years, coming out of the University of Manchester. jitLLM is essentially the LLM application layer built on top: the loop is Java, TornadoVM turns the hot matrix operations into CUDA kernels, and it can also bind existing CUDA libraries directly from Java code.
It grew out of GPULlama3.java. The engine started as GPULlama3.java, a GPU-accelerated fork of the pure-Java Llama3.java project, again from UNIMAN. It reads standard GGUF files, the same format llama.cpp uses, and falls back to a CPU path when no GPU is present.
The 90% claim needs its exact scope. The number circulating on Reddit and aggregator sites is that jitLLM reaches about 90% of llama.cpp performance for local inference on NVIDIA GPUs, demonstrated on an RTX 5090 in the project's own recordings. Two honest caveats. First, this is the project's claim, not an independent benchmark. Second, llama.cpp is a heavily hand-tuned C++/CUDA codebase with years of kernel optimization. Getting within 10% of it from compiled Java, if it holds across hardware and quantization formats, would be genuinely remarkable. Treat it as promising until someone replicates it on a plain RTX 4060 with a Q4 model, which is what most of us actually run.
What "Java compiled to CUDA" means in practice
The usual objection is that Java cannot touch a GPU without JNI and native libraries. TornadoVM sidesteps that with a different execution model.
- Annotated bytecode becomes kernels. You mark methods for GPU execution, and TornadoVM's JIT compiler lowers them to the target backend. On NVIDIA with the newer toolchains, that includes generating CUDA C and cuTile code at runtime.
- Memory is managed, not malloc'd. TornadoVM reuses Java memory where possible and handles device transfers, which removes most of the boilerplate you would write by hand against the CUDA API.
- Backend choice is a runtime property. The same code can target CUDA, OpenCL, or SPIR-V, so it is not locked to NVIDIA hardware. Apple Silicon and other OpenCL-capable devices are in scope.
If you have followed the Java Vector API and Panama work, this is the same family of ideas taken further: stop treating the JVM as a walled garden and let its bytecode be compiled to whatever hardware is underneath.
The part that makes it interesting for working teams
Raw performance is the headline, but the integration story is why I think this one is worth watching.
It is an official LangChain4j inference engine. You add langchain4j-jitllm and construct a JitLLMChatModel with onGPU(true), and every LangChain4j AI Service, tool-calling agent, and retrieval pipeline now runs on your own GPU. The ChatLanguageModel abstraction means the rest of your code does not know or care where inference happens.
There is a Quarkus extension too. quarkus-langchain4j-jitllm exposes the model as a CDI bean. The TornadoVM team demonstrated a standard Quarkus REST resource doing GPU inference in pure Java, built as part of the EU-funded AERO project with Red Hat. In their walkthrough, the entire wiring is a Quarkus property selecting jitLLM as the inference engine, plus a REST resource that injects the ChatModel bean. From the developer's side it looks identical to calling OpenAI or Ollama. That sameness is the point: the underlying article on tornadovm.org shows the full chain, Quarkus to LangChain4j to GPULlama3 to the GPU, with no code anywhere that mentions CUDA.
One JVM process for the whole stack. No Python interpreter, no model server process, no HTTP hop, no second logging and monitoring setup. For teams in healthcare, finance, or anywhere with data residency rules, "the model runs inside the same process as the application, on our hardware" is a much easier conversation than a data-flow diagram with an external API in the middle.
Context: how close is llama.cpp to the native ceiling anyway?
The 90% number only means something relative to a baseline, so it helps to know where llama.cpp itself sits. A widely read GitHub benchmark from the llama.cpp maintainers' discussion area compared llama.cpp against vLLM, the Python-native serving standard, on a single RTX 4090 serving a 3B model. For single parallel requests, llama.cpp needed between 93.6% and 100.2% of vLLM's runtime, and the two traded blows depending on context depth. In other words, for single-stream inference, llama.cpp is already close to what a purpose-built serving engine achieves. So "90% of llama.cpp" is really a claim of being within striking distance of the practical performance ceiling for this workload class, from Java bytecode, without writing a single CUDA kernel by hand.
It is also worth knowing how long this has been coming. The Quarkus integration was demonstrated back in June 2026 on the TornadoVM blog, built as part of the EU-funded AERO project, a collaboration between the University of Manchester and Red Hat. The October buzz is not a sudden appearance; it is the point where the earlier research demo matured into a published 1.0.0 toolkit with Maven artifacts and official framework integrations. That progression, research project to integration to standalone toolkit, is the normal shape of things getting real in the JVM ecosystem.
When you should NOT use it
Skepticism is warranted, and the project's own positioning invites some of it.
- It is early. Version 1.0.0, a young project, and the LangChain4j integration is still a beta artifact. Production teams should watch, prototype, and wait for independent benchmarks rather than bet a product launch on it.
- llama.cpp remains the safe default for raw local inference. If your architecture is already "llama.cpp server + a thin JVM client," that works, is battle-tested, and nothing here forces you to change. The aggregation and quantization ecosystem around llama.cpp is years ahead.
- Batch serving at scale is vLLM territory. The r/LocalLLaMA framing called it "Java vLLM-like," but vLLM's strengths, continuous batching and high-concurrency throughput, are a different problem than single-stream local inference. jitLLM competes in the laptop-and-single-server niche, not the inference-farm niche.
- NVIDIA-first today. CUDA and cuTile are the demonstrated path. The OpenCL and SPIR-V backends exist, but the claimed performance numbers are NVIDIA numbers.
A decision checklist for local inference on the JVM
Here is the takeaway I would actually save, as of October 2026:
-
Prototype or hobby project, want GPU speed in Java: jitLLM is the most direct path that exists. Add the Maven artifact, point it at a GGUF file, flip
onGPU(true). -
Production Spring Boot or Quarkus app, data cannot leave your servers: LangChain4j behind either jitLLM (if you accept early-stage risk) or Ollama/llama.cpp as the provider. Either way, keep the
ChatModelabstraction so you can swap later. - High-concurrency serving, many parallel sessions: vLLM or an inference farm. Java-native inference is not the right tool for this job yet.
- Just experimenting with local LLMs: llama.cpp with a UI on top. Simplest thing that works.
The deeper point is bigger than this one project. For a decade the answer to "GPU compute from Java" was "call native code." TornadoVM and jitLLM are the strongest evidence yet that the answer is changing to "let the JVM compiler target the GPU directly." If a 25-year-old managed runtime can get within 10% of hand-written CUDA, the sidecar era of Java AI may be shorter than we assumed.
I write about Java, Spring Boot, and AI engineering every week. Subscribe, it is free.
Have you tried jitLLM or GPULlama3.java on your own hardware? What tokens-per-second did you see, and on what GPU? I am collecting real numbers before I trust that 90% claim, and I would genuinely like to hear what you found.
Sources:
Top comments (0)