Shadows: High or Medium?
Textures: Will lowering them help, or just make everything look worse?
View distance: Is the graphics card struggling, or is the processor the bottleneck?
My friend struggles with finding the balance between frame rate and visual quality. That became the starting point for FrameForge, an offline Windows application that tests graphics configurations on the player’s own PC.
I could have built another tool that says, “Here are the best settings.”
Instead, I built one that asks:
Which settings are actually worth changing on this machine—and can we measure enough evidence to recommend them?
FrameForge uses TabPFN, an open-weight model for tabular data, to choose promising configurations. It then runs real tests using Intel PresentMon.
The distinction matters:
The model is allowed to guess which settings we should test next. It is not allowed to turn that guess into the final FPS result.
The final numbers come from measurements.
And if the improvement is too small to distinguish from ordinary run-to-run variation, the useful answer is not a victory banner.
It is: keep your settings.
What I Built
FrameForge is a local, adaptive game-settings optimizer built around a measurement loop.
You choose a supported game, a performance goal, and a minimum estimated visual-quality level. FrameForge changes the game’s configuration between runs, measures performance in a scene you select, and learns from the results.
It is not a chatbot. There is no prompt asking an LLM to invent “optimal settings.”
Its input is a small table:
Graphics settings → measured average FPS and 1% lows.
Its job is to decide which row would be most useful to measure next.
TL;DR
- Built for a friend: less settings-menu guesswork, more time actually playing.
- Real measurements: Intel PresentMon captures frame timing through Windows Event Tracing.
- Small-data AI: TabPFN guides the search using measurements collected during the session.
- Smoothness matters: the objective includes both average FPS and 1% lows.
- Player-controlled trade-off: candidates must satisfy the selected estimated quality floor.
- Local inference: bundled model weights, no cloud inference service.
- Reversible changes: configuration backups and recovery are part of the workflow.
- Measured confirmation: the selected configuration and baseline are retested before the final comparison.
FrameForge’s desktop interface. The workflow begins with the game and the player’s goal—not a chat box.
The Problem Isn’t Finding a Preset. It’s Finding the Right Compromise.
“Turn everything down” is an easy recommendation.
It is also a poor answer to someone who wants the game to look good.
Different settings have different costs. Their effects also depend on the hardware, the game, and the scene. Lowering something visually important but computationally cheap can produce a worse-looking game without solving the bottleneck.
A preset cannot settle that question for every PC.
Neither can a guide recorded on somebody else’s hardware.
| Approach | What it gives you | What remains unanswered |
|---|---|---|
| Low / Medium / High presets | A convenient starting point | Which compromises matter on your PC? |
| A settings guide | Someone else’s tested recommendations | Do those results transfer to your hardware and scene? |
| Manual tweaking | Direct control | How do you compare combinations fairly without spending the evening testing? |
| An FPS counter | A performance number | What should you change next? |
| FrameForge | A guided series of local experiments | Its answer still depends on the scene and measurements you give it |
That last limitation is deliberate.
FrameForge does not know every moment of a game. It knows the experiments performed in the selected scene.
That is a narrower claim than “perfect settings.”
It is also a more useful one.
Demo: From “Maybe Lower Shadows” to a Measured Decision
The session has four parts.
1. Choose what you care about
Pick a supported game and a goal: Smoothest, Balanced, or Best looking.
Then choose the minimum estimated picture quality you want to preserve.
The word estimated matters. FrameForge’s quality score comes from weights assigned to settings. It is not a measurement of your screen, and “85% quality” does not mean your eyes will perceive exactly 85% of the original image quality.
It is a practical constraint for the search—not an objective truth about beauty.
2. Measure the starting point
FrameForge launches the game.
You return to a repeatable scene and press F9 when ready. PresentMon records frame timing during the measurement window.
This establishes a baseline using the settings already on your machine.
3. Let the tuner choose the next experiment
FrameForge changes settings between runs and restarts the game.
You return to the same scene. Another measurement becomes another row in the session’s table.
The interface supports test budgets of 8, 12, or 20 runs, with early stopping when the tuner estimates that further improvement is unlikely.
This is not a one-click background miracle. You still need to reproduce the scene consistently.
The application handles configuration changes, measurement, and comparison. You provide the repeatable gameplay.
4. Check the result before accepting it
The selected configuration and the original baseline are measured again.
You see the comparison and the changed settings, then decide whether to use them.
A prediction selects an experiment. A measurement supports a recommendation.
The optimizer works through a bounded testing session rather than silently applying a generic preset.
Project: FrameForge landing page
Source: GitHub repository
How I Built It
FrameForge has three main components:
Desktop interface
│
▼
Rust / Tauri core
├── Game launch and session management
├── Configuration backup and modification
├── PresentMon frame capture
└── Python sidecar communication
│
▼
TabPFN tuner
Choose the next configuration
The interface uses vanilla HTML, CSS, and JavaScript inside Tauri. The native core is written in Rust. The AI engine is a Python sidecar packaged with PyInstaller.
Rust and Python communicate through line-delimited JSON over stdin and stdout.
That keeps the inference engine local without requiring a hosted API or a separate HTTP service.
The interesting part, though, is not the language split.
It is how the application decides what to measure.
Why TabPFN Fits This Problem
A gaming benchmark is an expensive way to create one training row.
Every observation involves changing settings, loading the game, reaching the scene, and running the measurement.
That means FrameForge cannot assume it will have thousands of examples from one player’s session.
It needs to make decisions from a small table.
TabPFN is designed for tabular prediction with small datasets. FrameForge uses it as a surrogate model: a model that approximates the expensive function we actually care about.
Here, that function is:
Given these graphics settings, what performance will this PC produce in this scene?
The surrogate is not the authority.
It helps spend the next benchmark run wisely.
The search does not begin by pretending one row is enough
The engine starts with the current configuration.
Its next suggestion uses the adapter’s prior estimates of performance cost and visual quality. With fewer than four observations, it explores different allowed configurations rather than relying on a premature model fit.
Once enough observations exist, TabPFN guides the search.
That separation is important:
Priors help us begin. Measurements give us evidence.
What the model sees
A configuration is represented using:
- normalized graphics-setting levels;
- an estimated performance-cost feature;
- an estimated visual-quality feature.
The measured targets are average FPS and 1% low FPS.
The engine proposes a pool of up to approximately 4,000 candidates, mixing random configurations with mutations around the best configuration found so far.
It removes duplicates, configurations already tested, and candidates below the selected quality floor.
TabPFN then predicts performance for the remaining candidates.
Why uncertainty matters
The engine requests the 16th, 50th, and 84th percentiles of each prediction.
It uses the median as a central estimate and the spread between quantiles to approximate uncertainty.
An Expected Improvement calculation then balances two reasons to test something:
- It looks likely to perform better.
- Its uncertainty makes it worth investigating.
Always choosing the highest predicted FPS can get stuck exploiting an inaccurate model.
Testing only uncertain configurations can waste the player’s time.
The acquisition function is the compromise.
There is an important caveat: FrameForge’s displayed “chance of improvement” comes from this approximation. It should not be treated as a thoroughly calibrated probability for every game and PC.
A High Average Can Still Feel Bad
Average FPS is useful.
It is not the whole experience.
A run can have a healthy average while its slower frames produce noticeable hitches. FrameForge therefore includes 1% lows alongside average performance.
The implemented search objective is:
score =
min(average FPS, target refresh rate)
+ 0.5 × min(1% low FPS, target refresh rate)
The visual-quality floor is handled separately as a constraint.
This encodes a particular preference: reaching the target and improving slower frames matter more than endlessly increasing the average beyond that target.
It is not a universal rule. Competitive players may value FPS above display refresh for latency reasons.
But an optimizer has to make its priorities explicit.
Otherwise, “best” is just a word hiding a trade-off.
The Most Important Engineering Work Wasn’t the Model
An optimizer that edits a friend’s game settings has responsibilities beyond prediction accuracy.
Back up before changing anything
FrameForge copies the configuration file before testing.
The recovery workflow restores the original settings after a stopped or failed session. If the application or PC shuts down mid-test, recovery is designed to run on the next launch.
The principle is simple:
A cancelled experiment should not become a permanent surprise.
Check that the game accepted the settings
Writing a file does not prove the game used the intended configuration.
FrameForge rereads the configuration after a run. If the game reverted values, that experiment is not counted as valid evidence.
Repeated application failures stop the session instead of allowing misleading comparisons to accumulate.
Watch temperature where supported
The current temperature guard supports NVIDIA GPUs.
AMD and Intel temperature monitoring are not implemented yet. That is a limitation to show clearly, not hide behind a generic “hardware safety” claim.
Measure without injecting into the game
PresentMon uses Windows Event Tracing to capture presentation timing. FrameForge does not need to inject code into the game or read its memory.
That does not establish compatibility with every anti-cheat system.
Testing belongs in offline, practice, or workshop modes, not ranked or live online matches.
Retest before declaring a winner
A promising run might simply have been an easier run.
FrameForge repeats measurements of the selected configuration and baseline to reduce the chance of recommending a noise-sized improvement.
Two runs are not a comprehensive statistical study. But they are a better basis for a recommendation than one lucky number.
What the Benchmarks Actually Prove—and What They Don’t
The repository includes a stress-test harness for evaluating surrogate models and search strategies.
It explores questions including:
- prediction accuracy with small sample sizes;
- sensitivity to noisy measurements;
- uncertainty coverage;
- ranking configurations by 1% lows;
- behavior with missing inputs;
- inference cost.
The current harness runs on a synthetic frame-time simulator.
That is useful for investigating algorithms under controlled conditions.
It is not evidence that FrameForge improves Counter-Strike 2 or Palworld by a particular percentage.
Those are different claims:
| Evidence | What it can support |
|---|---|
| A real before/after session | An observed improvement on that PC, in that scene |
| Repeated sessions across games and hardware | A broader claim about reliability |
| A friend’s actual use and feedback | Whether the workflow solves the intended person’s problem |
A good surrogate model is not automatically the best end-to-end optimizer.
That is why the project needs both algorithm evaluation and real gameplay measurements.
Why Does Open Innovation Matter?
FrameForge’s model runs on the same PC as the game.
The engine bundles its TabPFN checkpoint, enables offline settings, disables model telemetry, and blocks Python connection attempts through the socket functions it overrides.
That is a concrete local-inference design—not just a privacy paragraph attached to a cloud API.
It is also not a complete, independently audited proof of every possible network path. A recorded session with networking disabled and application-level request monitoring would make the offline claim easier for other people to verify.
Open weights make this design possible.
They let me build a tool that does not need:
- an inference account;
- an API key;
- a token budget;
- a hosted model to remain available.
The surrounding open-source tools matter too:
- Tauri provides the desktop foundation.
- PresentMon provides inspectable frame-timing infrastructure.
- Game adapters describe supported configuration files and settings.
Open innovation is useful here because another developer can inspect the measurement path, challenge the search strategy, or add support for another game.
They do not have to trust a screenshot that says “optimized.”
Limitations I Want Visible
FrameForge is an early implementation, not a universal performance fix.
- Game support is limited. Counter-Strike 2 and Palworld Steam versions have early support and all the unreal engine games only.
- Results are scene-dependent. A configuration that helps one area may behave differently elsewhere.
- Repeatability depends partly on the player. Different movement, weather, enemies, or background activity can distort comparisons.
- Visual quality is a heuristic. The application does not measure perceptual similarity from rendered images.
- CPU bottlenecks may leave little room for improvement. Lower graphics settings cannot solve every performance problem.
- Local inference takes resources. CPU latency is part of the user experience, not an implementation detail to ignore.
- Temperature monitoring currently covers NVIDIA only.
- Repeated confirmation reduces uncertainty; it does not eliminate it.
These are not reasons to abandon the idea.
App was tested on Black Myth Wukong - benchmark tool and showed clear improvement on it.
What I’d Build Next
Better estimates of measurement noise
Use repeated observations to distinguish a genuinely better configuration from normal variation more reliably.
Better quality preferences
A single quality slider is convenient, but players care about different things.
Future constraints could include “keep textures unchanged” or “never reduce view distance.”
More adapters
Expand game support through small, inspectable descriptions of configuration paths, values, and settings behavior.
Feedback from the person it was built for (playing CS2)
“Hold on. You did all that on its own?”
“Yeah. I usually tweak a few settings, start the game, play for a bit, look at the FPS, then adjust something else. After that I restart and do it again. You basically handled the tedious steps.”
“that helps”
"Time saver fr"
Code
The source, architecture, and build instructions are available in the FrameForge repository.
The landing page explains the testing workflow, safety protections, and current limitations.
Prize Categories
TabPFN Challenge
TabPFN is central to FrameForge’s adaptive search. It uses the session’s small table of measured configurations to predict candidate performance and help choose the next experiment.
This is not AI added to decorate the interface.
It changes how the application spends its limited testing budget.
The Point Was Never to Win an FPS Screenshot
A friend should not need to understand Bayesian optimization to choose between shadows and smoothness.
That is my job as the developer: put the complicated reasoning behind a workflow they can understand, inspect, and undo.
FrameForge cannot create hardware performance that is not there.
It can help answer a smaller, practical question:
Am I giving up image quality where it actually buys me something?
And it can insist on measuring before answering.
Let the AI choose the experiment. Let the PC provide the evidence. Let the player make the decision.







Top comments (1)
forgot to add these instructions: