<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: shanmukhamanidhar</title>
    <description>The latest articles on DEV Community by shanmukhamanidhar (@shanmukhamanidhar).</description>
    <link>https://dev.to/shanmukhamanidhar</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147506%2F24d4283d-cf14-4ee3-8253-71ccb8ec94a9.png</url>
      <title>DEV Community: shanmukhamanidhar</title>
      <link>https://dev.to/shanmukhamanidhar</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/shanmukhamanidhar"/>
    <language>en</language>
    <item>
      <title>SatQuery Al: Building a Conversational System for Satellite-Image Analysis</title>
      <dc:creator>shanmukhamanidhar</dc:creator>
      <pubDate>Mon, 28 Sep 2026 15:40:10 +0000</pubDate>
      <link>https://dev.to/shanmukhamanidhar/satquery-al-building-a-conversational-system-for-satellite-image-analysis-15be</link>
      <guid>https://dev.to/shanmukhamanidhar/satquery-al-building-a-conversational-system-for-satellite-image-analysis-15be</guid>
      <description>&lt;p&gt;SatQuery AI: Building a Conversational System for Satellite-Image Analysis&lt;br&gt;
I built SatQuery AI around a simple idea: I wanted users to interact with satellite and Earth-observation imagery through natural language rather than having to translate every question into a sequence of specialized image-processing and geospatial operations.&lt;br&gt;
A user should be able to ask:&lt;br&gt;
“Where has vegetation decreased?”&lt;br&gt;
or:&lt;br&gt;
“What changed between these two satellite images?”&lt;br&gt;
or:&lt;br&gt;
“Detect buildings in this region.”&lt;br&gt;
But building a system that answers these questions reliably is fundamentally different from building a chatbot.&lt;br&gt;
The distinction I kept coming back to was:&lt;br&gt;
A language model can explain an answer, but the satellite-analysis pipeline has to provide the evidence.&lt;br&gt;
That distinction shaped the architecture of SatQuery AI.&lt;br&gt;
From Questions to Evidence&lt;br&gt;
A conventional conversational AI system can take a question, generate an answer, and return it directly. That approach is not sufficient when the answer depends on information contained in an image.&lt;br&gt;
Suppose I ask:&lt;br&gt;
“Show me the areas where vegetation decreased between these images.”&lt;br&gt;
A language model can generate a plausible explanation of vegetation change. But generating a sentence is not the same as detecting vegetation change.&lt;br&gt;
For SatQuery AI, the workflow is closer to:&lt;br&gt;
Ask → Understand → Analyze → Verify → Visualize → Explain&lt;br&gt;
The natural-language layer first needs to understand what the user is asking. The system then needs to determine what type of Earth-observation analysis is appropriate and execute that analysis against the available imagery.&lt;br&gt;
Depending on the request, that could involve object detection, segmentation, change detection, image comparison, vegetation analysis, land-use and land-cover analysis, object counting, or other geospatial operations.&lt;br&gt;
The important part is that the analytical pipeline produces an actual result.&lt;br&gt;
For the vegetation example, that result might contain regions identified as changed, measurements associated with those regions, percentages, confidence information, or other relevant geospatial information.&lt;br&gt;
Only after that evidence exists does the conversational layer have something meaningful to explain.&lt;br&gt;
Separating Understanding from Execution&lt;br&gt;
One architectural decision I found important was keeping natural-language understanding separate from analytical execution.&lt;br&gt;
Conceptually, the system can be viewed as several layers:&lt;br&gt;
Natural-Language Query&lt;br&gt;
        ↓&lt;br&gt;
Query Understanding&lt;br&gt;
        ↓&lt;br&gt;
Analysis Planning&lt;br&gt;
        ↓&lt;br&gt;
Analytical Execution&lt;br&gt;
        ↓&lt;br&gt;
Evidence / Results&lt;br&gt;
        ↓&lt;br&gt;
Visualization&lt;br&gt;
        ↓&lt;br&gt;
Natural-Language Explanation&lt;br&gt;
The language model is therefore not treated as the source of truth for visual analysis.&lt;br&gt;
Its job is to understand the request and communicate the result. The specialized analytical pipeline is responsible for producing the underlying evidence.&lt;br&gt;
A simplified representation of the routing logic looks like this:&lt;br&gt;
query = "Show me the areas where vegetation decreased"&lt;/p&gt;

&lt;p&gt;intent = understand_query(query)&lt;/p&gt;

&lt;p&gt;if intent.type == "vegetation_change":&lt;br&gt;
    result = run_change_analysis(&lt;br&gt;
        image_a,&lt;br&gt;
        image_b&lt;br&gt;
    )&lt;/p&gt;

&lt;p&gt;answer = explain_result(result)&lt;br&gt;
The actual implementation can become considerably more complex, but the separation is important. It prevents the conversational component from becoming responsible for calculations and visual conclusions it cannot independently establish.&lt;br&gt;
Making the Result Visual&lt;br&gt;
Another important part of the system is that the result should not end as text.&lt;br&gt;
If an analysis identifies regions of vegetation change, those regions need to be represented on the imagery or map.&lt;br&gt;
Conceptually:&lt;br&gt;
result = analyze(images)&lt;/p&gt;

&lt;p&gt;visual_layer = create_visualization(&lt;br&gt;
    image=images,&lt;br&gt;
    regions=result.regions&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;return {&lt;br&gt;
    "evidence": result,&lt;br&gt;
    "visualization": visual_layer&lt;br&gt;
}&lt;br&gt;
This gives the user two complementary forms of information.&lt;br&gt;
The visualization answers:&lt;br&gt;
“Where did this happen?”&lt;br&gt;
The analytical result answers:&lt;br&gt;
“What did the system measure?”&lt;br&gt;
And the natural-language explanation answers:&lt;br&gt;
“What does this result mean?”&lt;br&gt;
Keeping these three responsibilities distinct makes the interaction easier to reason about.&lt;br&gt;
Before and After: Turning a Query into an Analysis&lt;br&gt;
Without this separation, a conversation might look like this:&lt;br&gt;
User&lt;br&gt;
Show me the areas where vegetation decreased between these images.&lt;br&gt;
AI&lt;br&gt;
Vegetation decreased in several areas between the two images.&lt;br&gt;
That answer sounds reasonable, but it provides little evidence.&lt;br&gt;
With the SatQuery approach, the interaction is intended to become:&lt;br&gt;
User&lt;br&gt;
Show me the areas where vegetation decreased between these images.&lt;br&gt;
System&lt;br&gt;
Identifies the request as a vegetation-change analysis.&lt;br&gt;
Analysis pipeline&lt;br&gt;
Processes the two images and identifies relevant regions.&lt;br&gt;
Visualization&lt;br&gt;
Highlights those regions on the satellite imagery.&lt;br&gt;
System&lt;br&gt;
Reports the resulting measurements and relevant confidence information.&lt;br&gt;
AI&lt;br&gt;
Explains what the analysis found in natural language.&lt;br&gt;
The difference is subtle from the user's perspective, but significant from an engineering perspective. The second workflow has an explicit analytical stage between the question and the answer.&lt;br&gt;
Adding Conversational Memory with Hindsight&lt;br&gt;
Once the system can perform individual analyses, another problem appears: users naturally want to continue the conversation.&lt;br&gt;
For example:&lt;br&gt;
User&lt;br&gt;
Show me the areas where vegetation decreased between these images.&lt;br&gt;
After receiving the result, the user might ask:&lt;br&gt;
Now compare those regions with the previous analysis.&lt;br&gt;
The second question depends on context from the first interaction.&lt;br&gt;
I integrated Hindsight as the agent-memory layer for this part of SatQuery AI. Hindsight is designed to provide persistent memory for AI agents, allowing useful information from previous interactions to remain available instead of treating every interaction as completely isolated.&lt;br&gt;
Hindsight on GitHub⁠�&lt;br&gt;
Hindsight Documentation⁠�&lt;br&gt;
Vectorize — What is Agent Memory?⁠�&lt;br&gt;
The distinction I found useful is that memory should preserve context, while the analytical pipeline should preserve evidence.&lt;br&gt;
For example, memory can help establish that “those regions” refers to regions identified during an earlier vegetation analysis. It does not mean that memory itself becomes the source of the satellite-analysis result.&lt;br&gt;
Conceptually:&lt;br&gt;
context = hindsight.retrieve(query)&lt;/p&gt;

&lt;p&gt;analysis_request = understand_query(&lt;br&gt;
    query=query,&lt;br&gt;
    context=context&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;result = execute_analysis(analysis_request)&lt;/p&gt;

&lt;p&gt;hindsight.remember(&lt;br&gt;
    query=query,&lt;br&gt;
    result_context=result&lt;br&gt;
)&lt;br&gt;
This makes the interaction conversational without collapsing the boundaries between memory, reasoning, and analysis.&lt;br&gt;
The Architecture I Ended Up Thinking About&lt;br&gt;
I think about SatQuery AI as six cooperating components rather than one large AI system.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Natural-Language Understanding
This layer interprets what the user wants.
It extracts the intent and relevant parameters needed to construct an analytical request.&lt;/li&gt;
&lt;li&gt;Analytical Execution
This is where the actual image and geospatial processing happens.
The appropriate operation is executed against the supplied Earth-observation data.&lt;/li&gt;
&lt;li&gt;Evidence and Results
The output of the analytical pipeline is treated as the evidence layer.
This can include detected regions, object counts, change areas, percentages, confidence information, and geospatial information.&lt;/li&gt;
&lt;li&gt;Visualization
The analytical results are translated into visual overlays or other representations that can be displayed on satellite imagery or a map.&lt;/li&gt;
&lt;li&gt;Conversational Memory
Hindsight provides persistent agent memory so that subsequent interactions can refer to relevant previous context.&lt;/li&gt;
&lt;li&gt;Natural-Language Explanation
Finally, the conversational layer turns the analytical result into an explanation that a user can understand.
This separation gives each component a relatively clear responsibility.
A Design Challenge: Memory Is Not Ground Truth
One of the challenges I had to keep in mind is that persistent memory can make an agent feel more intelligent while also creating another source of potential confusion.
If a previous interaction says that a particular region was important, a later query may depend on that context. But remembered context should not automatically be treated as a fresh analytical conclusion.
This is especially important for Earth-observation systems because the underlying data can change, the user can switch datasets, and the meaning of a query can depend on the imagery being analyzed.
I therefore see memory as a mechanism for continuity, not a replacement for analytical verification.
The same principle applies to the language model itself.
A model can interpret:
“Compare those regions with the previous analysis.”
But it should not invent what “those regions” contain. The system needs to connect that reference to previously established analytical context and, where necessary, perform the appropriate analysis again.
What I Learned Building It&lt;/li&gt;
&lt;li&gt;AI reasoning and evidence should be separate
A language model is very good at interpreting language and explaining information. That does not make it an image-analysis engine. Keeping those responsibilities separate makes the system easier to reason about.&lt;/li&gt;
&lt;li&gt;Visualization is part of the answer
For geospatial problems, coordinates, regions, overlays, and measurements can communicate information that text alone cannot. A textual answer without a corresponding visual result can leave an important part of the user's question unanswered.&lt;/li&gt;
&lt;li&gt;Conversational memory needs boundaries
Persistent memory makes multi-step interactions much more natural, but remembered context should not be confused with verified analytical evidence.&lt;/li&gt;
&lt;li&gt;The analytical pipeline should remain modular
Object detection, segmentation, change detection, vegetation analysis, and other operations should be treated as capabilities that can be composed rather than embedding every operation directly inside the conversational layer.&lt;/li&gt;
&lt;li&gt;The hardest part is connecting the layers
The interesting engineering problem is not simply adding an LLM to a satellite-image application. It is designing the interfaces between language understanding, analytical execution, evidence, visualization, memory, and explanation.
Conclusion
SatQuery AI changed how I think about conversational interfaces for visual and geospatial systems.
The goal is not to make a language model pretend that it can answer every question about an image. Instead, the language model becomes the interface through which a user describes an analytical task.
The underlying Earth-observation pipeline produces the evidence.
The visualization makes that evidence inspectable.
The memory layer preserves conversational continuity.
And the language model explains the resulting information.
That separation leads back to the principle that shaped the project:
A language model can explain an answer, but the satellite-analysis pipeline has to provide the evidence.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
