DEV Community

Cover image for A beginner's guide to the Lotus-Diffusion-Dense-Prediction model by Kjjk10 on Replicate
aimodels-fyi
aimodels-fyi

Posted on Originally published at aimodels.fyi

A beginner's guide to the Lotus-Diffusion-Dense-Prediction model by Kjjk10 on Replicate

This is a simplified guide to an AI model called Lotus-Diffusion-Dense-Prediction maintained by Kjjk10. If you like these kinds of analysis, you should join AImodels.fyi or follow us on Twitter.

Overview

lotus-diffusion-dense-prediction estimates depth from an input image and returns a depth visualization or grayscale depth map. The Replicate listing is maintained by kjjk10. Lotus is a diffusion-based visual foundation model adapted for dense prediction: instead of predicting noise through a multi-step diffusion process, it predicts annotations in a single step. The paper describes this design as a way to avoid harmful variance, simplify optimization, and improve inference speed; it also introduces a “detail preserver” tuning strategy for finer predictions. The authors report zero-shot depth and surface-normal estimation results across datasets, but this Replicate endpoint exposes depth estimation only. The most important practical point is that the API’s disparity option changes the color mapping for visualization; it does not turn the endpoint into a separate depth-estimation task. The supplied materials do not state parameter count, exact processing resolution, latency, or per-prediction cost.

Best use cases

Depth-map generation for image understanding prototypes. Provide a photograph and use the returned map to test whether a downstream system can use scene layout, relative distance, or object boundaries. Lotus was designed for zero-shot dense geometry prediction, so it fits experiments that need a depth estimate without task-specific training.

Depth-guided compositing and editing. Use the grayscale output as a structural guide for separating foreground and background or for testing depth-aware effects. The endpoint returns a URI for the result, and the schema lets you choose grayscale output rather than a colorized visualization.

Single-image 3D reconstruction experiments. Use the predicted depth map as an input to a reconstruction pipeline or as a way to inspect whether a scene’s broad geometry is plausible. The paper identifies single- and multi-view 3D reconstruction as practical applications, but this endpoint itself returns a depth map, not a point cloud or 3D asset.

Benchmarking zero-shot depth estimation. Researchers can run images through the endpoint and compare its predictions with reference depth maps. The project README describes evaluation on depth datasets and provides evaluation scripts for the repository, but the Replicate API schema does not expose benchmark metrics or evaluation controls.

Limitations

  • Depth only in this endpoint. Lotus also supports surface-normal estimation in the project, but the supplied Replicate schema describes an input image for depth estimation and provides no task selector for normals.
  • No metric-depth guarantee is documented. The schema describes a colorized depth map or “raw depth” grayscale output, but it does not specify units, calibration, scale, or whether values correspond to metric distances. Do not treat the output as calibrated sensor depth without validation.
  • No documented image-size limits. processing_resolution controls the maximum resolution used for processing, but the schema gives no allowed values, upper bound, or exact default resolution beyond 0. The README does not state output dimensions or supported image formats.
  • No latency, cost, or hardware figures for this endpoint. The paper claims Lotus is faster than most existing diffusion-based methods, but the supplied materials give no measured runtime for this Replicate deployment. The repository’s tested setup used Ubuntu 20.04 LTS, Python 3.10, CUDA 12.3, and an NVIDIA A800-SXM4-80GB; that is a tested development environment, not a stated Replicate requirement.
  • Prediction quality still needs task-specific checks. The sources report strong zero-shot results, but provide no failure-case list or guarantee for unusual scenes, transparent surfaces, reflective materials, low light, or fine structures. Validate outputs on the image types and downstream decisions that matter to your application.
  • The endpoint’s disparity description is easy to misread. It says the option reverses the color mapping for disparity visualization. It does not document a change to the underlying prediction or a separate disparity output.
  • License: the metadata points to the project’s Apache-2.0 license. Review the repository license and any applicable deployment terms before commercial use.
  • Model status: the Replicate metadata records a latest version created on 2025-01-09. The README notes newer Lotus model releases for some tasks, including updated normal models and disparity variants, but the supplied information does not establish whether this specific Replicate endpoint receives ongoing updates.

How it compares

  • lotus is the closest alternative: it is described as the same diffusion-based visual foundation model for dense prediction. Choose lotus-diffusion-dense-prediction when you need this endpoint’s documented controls for output type, resampling, processing resolution, and disparity visualization; choose lotus if its own interface or deployment better fits your workflow. The supplied information does not provide comparative latency, quality scores, or prices for the two Replicate endpoints.
  • stable-diffusion-depth2img creates image variations while preserving shape and depth. Choose Lotus when the job is to estimate depth from an image; choose stable-diffusion-depth2img when the goal is image generation or editing guided by depth structure. These are different tasks, so their outputs and quality should not be compared as if they were interchangeable depth estimators.
  • ssd-lora-inference is a proof of concept for inference on SSD-1B LoRAs. Choose Lotus for dense depth prediction; choose ssd-lora-inference when you need to run an SSD-1B LoRA. The supplied description does not give enough information to compare their speed, cost, or output quality.
  • midas is explicitly described as robust monocular depth estimation. Choose Lotus if you want to evaluate the diffusion-based, single-step dense-prediction approach and its output controls; choose midas as a direct alternative for monocular depth estimation. No benchmark results, runtime figures, or prices for these Replicate endpoints are provided here, so test both on your own images before selecting one.
  • flux-control-lora-depth/image-to-image uses a depth map as a control image to transfer structure into a generated image, with another initial image guiding color. Choose Lotus to estimate depth from an input image; choose flux-control-lora-depth/image-to-image when you already have a depth control image and want guided image generation. The key difference is prediction versus generation, not a documented speed or quality ranking.

Technical specifications

Lotus adapts visual priors from pretrained text-to-image diffusion models for dense prediction. The paper’s central changes are direct annotation prediction instead of noise prediction, a single-step reformulation instead of multi-step noising and denoising, and a detail-preserver tuning strategy. The authors report zero-shot depth and normal estimation without scaling up training data or model capacity. The supplied materials do not identify the backbone, parameter count, training-set size, model file format, quantization, or exact benchmark scores.

The repository README documents a tested environment of Ubuntu 20.04 LTS, Python 3.10, CUDA 12.3, and NVIDIA A800-SXM4-80GB. It uses Conda and pip install -r requirements.txt; training scripts use Accelerate, tested with version 0.29.3. The README describes Hypersim and Virtual KITTI data preparation for training, including RGB, depth, and normal-map workflows, but does not state the total number of training examples or the exact data mixture used for this Replicate version.

Replicate metadata lists the model as public, with Cog version 0.13.6. Its latest version ID is 26033a14ffa3b967b54daaf425adeba698e9fa67260422b9e41c4907ae0e45f5, created on 2025-01-09 at 12:23:15.039555Z. The metadata points to the project’s Apache-2.0 license.

Model inputs and outputs

Inputs

  • image — string URI; required by the schema; input image for depth estimation.
  • output_type — enum reference named output_type; default: "color". The description lists "color" for a colorized depth map and "grayscale" for raw depth. The schema does not provide the full enum definition.
  • resample_method — enum reference named resample_method; default: "bicubic". The description says this selects a resampling method for quality; the schema does not list the other enum values.
  • processing_resolution — enum reference named processing_resolution; default: 0. The description calls this the maximum processing resolution and says higher values use more memory for better quality. The schema does not provide the allowed values or bounds.
  • disparity — boolean; default: true. The description says it reverses the color mapping for disparity visualization.

Outputs

  • Output — string URI. The schema does not specify a file extension, MIME type, image dimensions, or whether the URI points to a color or grayscale file; those depend on the selected output settings.

Getting started

Install the Replicate Python client and set REPLICATE_API_TOKEN in your environment. Replace the placeholder image URI with a URI accessible to Replicate.

import replicate

output = replicate.run(
    "kjjk10/lotus-diffusion-dense-prediction",
    input={
        "image": "https://example.com/input.jpg",
        "output_type": "color",
        "resample_method": "bicubic",
        "processing_resolution": 0,
        "disparity": True,
    },
)

print(output)
Enter fullscreen mode Exit fullscreen mode

The schema defines the result as a URI. Retrieve that URI with your application’s HTTP client if you need to save or process the output file.

Frequently asked questions

Q: What inputs does this Replicate model require?

A: It requires an image URI. You can also set output type, resampling method, processing resolution, and the disparity visualization flag; the documented defaults are color, bicubic, 0, and true.

Q: What output format does it return?

A: The Replicate output schema returns a string URI. The schema does not specify the file extension, MIME type, or image dimensions.

Q: Can I use this model commercially, and what license applies?

A: The supplied metadata points to the project’s Apache-2.0 license. Check the repository license and applicable Replicate terms for your intended use.

Q: Does the output provide metric depth?

A: The schema describes a colorized depth map or grayscale “raw depth,” but does not specify units, calibration, or metric scale. Validate the output before using it as a measurement source.

Q: What does the disparity input do?

A: It is a boolean that defaults to true and is described as reversing the color mapping for disparity visualization. The schema does not say that it changes the underlying depth prediction.

Q: How does Lotus compare with MiDaS?

A: Both are presented as depth-estimation options: Lotus uses a diffusion-based, single-step dense-prediction approach, while midas is described as robust monocular depth estimation. The supplied information does not include comparable benchmark, speed, or cost figures.

Q: Is this endpoint suitable for production use?

A: It can support production prototypes or workflows that can tolerate uncalibrated depth estimates, but the supplied materials do not document latency, availability targets, image-size limits, or metric-depth guarantees. Test it on representative inputs and verify the output against your application’s requirements.

Q: Is this specific Replicate version still actively maintained?

A: The metadata records a latest version created on 2025-01-09. The README documents later Lotus project updates, but the supplied information does not confirm ongoing updates to this Replicate endpoint.

Click here to read the full guide to Lotus-Diffusion-Dense-Prediction

Top comments (0)