This is a simplified guide to an AI model called Sa2va-4b-Image maintained by Bytedance. If you like these kinds of analysis, you should join AImodels.fyi or follow us on Twitter.
Overview
sa2va-4b-image is ByteDance’s image model in the Sa2VA family, which combines SAM-2 segmentation with a multimodal large language model (MLLM) to connect natural-language instructions with image regions. It accepts an image and a text instruction, then returns a text response and an image URI. The research describes Sa2VA as a unified system for question answering, visual-prompt understanding, referring segmentation, and grounded conversation across images and videos; its LLM generates instruction tokens that guide SAM-2 to produce masks. The project reports support for InternVL2.5/3 and Qwen2.5-VL/Qwen3-VL backbones, but the Replicate listing does not specify which backbone this 4B image endpoint uses. The maintainer is ByteDance. The key practical caveat is that this endpoint’s published schema exposes only a single image and instruction, not video input or a separate mask output: do not assume it provides the full video workflow or a machine-readable segmentation mask just because the broader Sa2VA research supports those tasks.
Best use cases
Ask questions about image content. Provide an image and a question such as “What is the person holding?” or “Describe the objects on the table.” Sa2VA is designed for multimodal question answering and grounded understanding, and its paper reports question-answering performance comparable to Qwen2-VL and InternVL2.5 on the benchmarks it evaluated. The available schema does not specify a response format or guarantee that answers include coordinates.
Request a description of a particular object or region. Instructions such as “Describe the red bag” or “Which object is closest to the door?” fit the model’s combination of language understanding and visual grounding. This is useful for image review and interactive visual search, but the endpoint schema does not expose a bounding-box field or a structured region identifier.
Prototype referring segmentation workflows. Sa2VA’s research focus includes referring segmentation: identifying an object from a text expression and producing a mask. The Replicate output schema includes an image URI as well as a text response, which may support a visual result, but it does not document the image’s contents or provide a mask file field. Test the actual returned image before building a pipeline that depends on segmentation output.
Build image-grounded conversational interfaces. A developer can pass a user’s image and instruction to the endpoint and display the returned response alongside the returned image URI. This suits prototypes that need a single-turn image interaction. The schema does not describe conversation history, so multi-turn state management must be handled by the application, and the endpoint’s support for passing prior turns is not established.
Limitations
-
The endpoint is narrower than the research system. The paper describes image and video tasks, but this Replicate schema accepts one
imageURI and oneinstructionstring. It does not expose video, frame sequences, timestamps, or video-length controls. -
The output contract does not guarantee a segmentation mask. The schema requires a text
responseand optionally lists animgURI. It does not define a mask tensor, polygon, bounding box, class label, or confidence score. Treat segmentation as a capability of the model family, not a guaranteed structured output from this endpoint. - No image-size or instruction limits are published here. The schema gives no maximum dimensions, file-size limit, accepted image MIME types, context window, or instruction-length constraint. Check the deployed endpoint’s behavior with representative inputs.
- No performance or hardware figures are provided. The available materials do not state latency, throughput, VRAM requirements, or per-call cost. The 4B label identifies the model variant, but the Replicate metadata does not provide a parameter-count specification beyond that name.
- No endpoint-specific quality breakdown is available. The paper reports strong results across tasks, especially referring video object segmentation, but the supplied materials do not give benchmark scores for this exact Replicate image endpoint. Results on a research benchmark do not establish accuracy for a particular production dataset.
- The license link identifies Apache 2.0, but deployment terms still matter. The supplied license metadata points to Apache 2.0. Review the applicable model and service terms for your use case; the materials here do not describe data retention, privacy, or service-level guarantees.
-
The endpoint’s maintenance status is not established. Replicate metadata records a latest version created on 2025-02-22 and a Cog version of
0.13.8-dev+gdeaa413.d20250220. That is a version record, not a promise of ongoing updates or support.
How it compares
-
Sa2VA-4B: Choose
sa2va-4b-imagewhen you want a hosted Replicate endpoint with a simple image-and-instruction API. Choose Sa2VA-4B when you want the Hugging Face model listing and need to evaluate that distribution route. The supplied information does not provide a measured speed, quality, or cost comparison between them. -
sa2va-26b-image: Choose
sa2va-4b-imagewhen the 4B variant and its two-field input schema fit your task; the smaller variant name may be relevant when model size is a concern, but no latency or cost figures are provided. Choose sa2va-26b-image when you want to evaluate the 26B image variant. The available materials do not establish that the larger variant produces better results or quantify its speed and cost tradeoffs. -
Sa2VA-8B: Choose
sa2va-4b-imagefor the hosted Replicate interface; choose Sa2VA-8B when you want to assess the 8B Hugging Face variant. The supplied sources do not give comparable benchmark scores, inference times, or deployment requirements for these variants. -
sa2va-8b-image: Choose
sa2va-4b-imagewhen you want the 4B image endpoint; choose sa2va-8b-image when you want to test the 8B Replicate image endpoint. Both are presented as image variants, but the supplied information does not document differences in input shape, output shape, quality, speed, or price. -
sa2va/8b/image: Choose
sa2va-4b-imagewhen you want the Replicate-hosted 4B endpoint. Choose sa2va/8b/image when you want to evaluate the 8B model on fal.ai. The available sources do not support a direct comparison of cost, latency, or output quality across the two hosting platforms.
Technical specifications
Sa2VA is a unified architecture that combines SAM-2, a foundation video segmentation model, with an MLLM. It maps text, images, and video into a shared LLM token space; the LLM generates instruction tokens that guide SAM-2 in producing masks. The project README describes support for InternVL2.5/3 and Qwen2.5-VL/Qwen3-VL backbones. The Replicate listing identifies this endpoint as the 4B image variant, but does not specify its exact backbone, image resolution, context window, quantization, model file format, or hardware requirements.
The paper introduces Ref-SAV, an auto-labeled dataset with more than 72,000 object expressions in complex video scenes. The authors also manually validate 2,000 video objects in Ref-SAV for referring video object segmentation evaluation. These are research dataset details; the supplied materials do not say that this Replicate endpoint was trained exclusively on Ref-SAV or disclose its full training data.
Confirmed endpoint and project details:
- Task family: Image-to-text; Sa2VA research capabilities include question answering, visual-prompt understanding, referring segmentation, grounded conversation, and image/video chat.
-
Replicate inputs:
imageandinstruction, both strings. -
Replicate outputs: Required text field
response; optional URI fieldimg. - License metadata: Apache 2.0 license URL.
- Replicate visibility: Public.
- Latest version created: 2025-02-22.
-
Cog version:
0.13.8-dev+gdeaa413.d20250220. -
Repository environment: The README uses
uvwith a projectpyproject.tomlanduv.lock; it documentsuv sync --extra=latestanduv sync --extra=legacy. The legacy option is described for InternVL2.5 or earlier. - Not specified in the supplied endpoint materials: Maximum image dimensions or file size, accepted image MIME types, default values, enum values, minimum or maximum instruction length, output image dimensions, latency, price, VRAM, and context window.
Model inputs and outputs
Inputs
-
image— string, URI format. Described as the input image for segmentation. The schema does not specify accepted file types, image dimensions, size limits, or a default. -
instruction— string. Described as a text instruction for the model. The schema does not specify a default, enum, or length constraint.
Outputs
-
response— required string containing the model’s text response. The schema does not define a response format or guarantee structured segmentation data. -
img— optional string in URI format. The schema does not specify whether this is an original image, an annotated image, or a segmentation visualization, so inspect the returned asset before relying on its contents.
Getting started
Install the Replicate Python client and set REPLICATE_API_TOKEN in your environment. Replace the placeholder image URI with a URI accessible to the endpoint.
import os
import replicate
output = replicate.run(
"bytedance/sa2va-4b-image",
input={
"image": "https://example.com/your-image.jpg",
"instruction": "Describe the main objects in this image.",
},
)
print(output)
The schema defines the output as an object with a required response and an optional img URI. The client’s returned value may be represented as a mapping or another SDK-supported object; inspect it before accessing fields in application code.
Frequently asked questions
Q: What inputs does sa2va-4b-image require?
A: The Replicate schema defines an image string in URI format and an instruction string. It does not publish defaults, accepted image types, or size and length limits.
Q: What output format does this Replicate model return?
A: The output schema requires a text response and lists an optional img URI. It does not define a separate mask, bounding box, or other structured segmentation field.
Q: Can I use sa2va-4b-image commercially? What license applies?
A: The supplied license metadata points to Apache 2.0. Check the applicable license and service terms for your deployment; the provided materials do not describe data-retention or privacy terms.
Q: Does this endpoint accept video?
A: No video input appears in the published Replicate schema; it lists one image URI and one instruction. The broader Sa2VA research covers video tasks, but that does not establish video support for this endpoint.
Q: Does the model always return a segmentation mask?
A: The Sa2VA architecture is designed for dense grounded understanding and mask generation, but this endpoint’s schema guarantees only a text response and an optional image URI. It does not promise a machine-readable mask or specify what the returned image contains.
Q: How does it compare with the 8B Replicate image variant?
A: sa2va-8b-image is the 8B image variant, while this endpoint is identified as 4B. The supplied sources do not provide comparable quality, latency, or cost measurements, so test both on your own images and instructions.
Q: Is sa2va-4b-image suitable for production use?
A: It can be evaluated in an application through its image-and-instruction API, but the supplied materials do not state latency, availability guarantees, privacy terms, or endpoint-specific benchmark results. Validate output behavior and operational requirements before relying on it in production.
Q: Is the model still actively maintained?
A: The metadata records a public latest version created on 2025-02-22 and a Cog version from February 2025. Those records do not confirm an ongoing maintenance schedule.
Top comments (0)