DEV Community

Umair Bilal
Umair Bilal

Posted on Originally published at buildzn.com

How I Cut LLM Costs 90% with Multi-LLM Chatroom CLI Agents

This article was originally published on BuildZn.

Everyone talks about multi-agent systems, but nobody explains how expensive they get when you're hitting multiple LLM APIs 24/7. Figured it out the hard way. Spent months building FarahGPT and NexusOS, saw the API bills, and knew there had to be a better way for a multi llm chatroom cli agents setup. This isn't just theory; this is how we built a cost-effective system.

The Real Cost of a Multi-LLM Chatroom CLI Agents Setup

Look, the hype around LLMs is real. But so are the API bills. Want a multi-agent system where Claude, ChatGPT, Grok, and maybe even Gemini are all talking to each other? That's a token torrent. Even with discounts, these things add up, especially if you're running complex reasoning chains or a high-volume flutter multi agent system. We're talking hundreds, sometimes thousands of dollars a month just for inference. For a side project, or even a startup trying to conserve runway, that's brutal.

The ai-chatroom concept is cool, but scaling it with direct API calls? Your wallet will hate you. Honestly, I think the current LLM API pricing models are overengineered for simple agent-to-agent communication. We need to focus on llm agent cost reduction, and that's exactly what this approach delivers.

The 90% Cost Reduction Claim: Here's the Math

Let's break down the cost. A typical multi-LLM conversation involving 3 agents, each making 2-3 turns, easily hits 5,000-10,000 tokens. If you run that 100 times a day, across multiple users, you're looking at millions of tokens.

  • Claude 3 Opus: ~$75/1M input tokens, ~$150/1M output tokens.
  • GPT-4 Turbo: ~$10/1M input tokens, ~$30/1M output tokens.
  • Gemini 1.5 Pro: ~$7/1M input tokens, ~$21/1M output tokens.

A mixed agent system using these directly for 10M tokens a month can easily run you $500 - $1500+.

With nodejs cli automation via browser control, you're paying for server uptime and IP rotation (if needed), not per token. A single EC2 instance (c5.large or similar) running Puppeteer could manage multiple concurrent LLM sessions for ~$100-$200/month. That's a 90% reduction compared to the high end of direct API usage, even when accounting for browser overhead and potential failed runs. This is the ultimate no api key ai agents play.

The Browser Automation Gambit for No API Key AI Agents

Here's the thing — every major LLM has a web UI. ChatGPT, Claude, Grok, Gemini, Perplexity... they're all there. What if instead of hitting their APIs, we just automated a browser to type into their chat boxes, read their responses, and pass them along? That's the core idea behind this cost-saving multi llm chatroom cli agents setup.

We're essentially turning each LLM's web interface into its CLI. Node.js, combined with Puppeteer (or Playwright, same difference), becomes our nodejs cli automation engine. It's a headless browser farm orchestrated by your backend.

How it works:

  1. Node.js Orchestrator: Your backend manages multiple browser instances, one for each LLM agent.
  2. Browser Sessions: Each browser logs into an LLM's web UI. Think persistent cookies, maybe even a dedicated profile.
  3. Input: When Agent A needs to talk to Agent B, your Node.js orchestrator uses Puppeteer to find the chat input box in Agent B's browser, types the message, and hits enter.
  4. Output: Puppeteer then waits for the LLM's response to appear, scrapes the text from the UI, and passes it back to your orchestrator.
  5. Relay: The orchestrator then decides which agent gets the response next, repeating the input/output cycle.

This creates a true browser based llm chat where the "chat" isn't between humans, but between automated LLMs.

Building the Flutter Multi Agent System & Node.js Orchestrator

Alright, let's get into the guts. You need a Flutter frontend to display the conversation and let a human kick it off. The heavy lifting is all in Node.js.

Flutter UI: The Frontend Glimpse

For the Flutter side, it's a standard chat interface. You'll send a prompt to your Node.js backend via an HTTP POST request or WebSockets, and the backend will stream back messages from the multi llm chatroom cli agents.

// lib/screens/chat_screen.dart
import 'package:flutter/material.dart';
import 'package:http/http.dart' as http;
import 'dart:convert';

class ChatScreen extends StatefulWidget {
  @override
  _ChatScreenState createState() => _ChatScreenState();
}

class _ChatScreenState extends State<ChatScreen> {
  final List<Map<String, String>> _messages = [];
  final TextEditingController _controller = TextEditingController();

  Future<void> _sendMessage(String text) async {
    setState(() {
      _messages.add({'sender': 'User', 'message': text});
    });
    _controller.clear();

    try {
      final response = await http.post(
        Uri.parse('http://localhost:3000/chat'), // Your Node.js backend
        headers: {'Content-Type': 'application/json'},
        body: json.encode({'prompt': text}),
      );

      if (response.statusCode == 200) {
        final data = json.decode(response.body);
        // Assuming your Node.js backend returns agent responses directly
        // For a true streaming setup, you'd use WebSockets.
        setState(() {
          _messages.add({'sender': data['agent'], 'message': data['response']});
        });
      } else {
        setState(() {
          _messages.add({'sender': 'System', 'message': 'Error: ${response.statusCode}'});
        });
      }
    } catch (e) {
      setState(() {
        _messages.add({'sender': 'System', 'message': 'Network Error: $e'});
      });
    }
  }

  @override
  Widget build(BuildContext context) {
    return Scaffold(
      appBar: AppBar(title: Text('CLI-Powered LLM Chatroom')),
      body: Column(
        children: [
          Expanded(
            child: ListView.builder(
              itemCount: _messages.length,
              itemBuilder: (context, index) {
                final msg = _messages[index];
                return Align(
                  alignment: msg['sender'] == 'User' ? Alignment.centerRight : Alignment.centerLeft,
                  child: Container(
                    padding: EdgeInsets.all(8),
                    margin: EdgeInsets.symmetric(vertical: 4, horizontal: 8),
                    decoration: BoxDecoration(
                      color: msg['sender'] == 'User' ? Colors.blue[100] : Colors.grey[200],
                      borderRadius: BorderRadius.circular(12),
                    ),
                    child: Text('${msg['sender']}: ${msg['message']}'),
                  ),
                );
              },
            ),
          ),
          Padding(
            padding: const EdgeInsets.all(8.0),
            child: Row(
              children: [
                Expanded(
                  child: TextField(
                    controller: _controller,
                    decoration: InputDecoration(hintText: 'Message LLM agents...'),
                    onSubmitted: _sendMessage,
                  ),
                ),
                IconButton(
                  icon: Icon(Icons.send),
                  onPressed: () => _sendMessage(_controller.text),
                ),
              ],
            ),
          ),
        ],
      ),
    );
  }
}
Enter fullscreen mode Exit fullscreen mode

This is basic. For a real flutter multi agent system chat, you'd use WebSockets for real-time updates from multiple agents, showing who's "typing" and receiving messages in sequence.

Node.js Backend: The Orchestration Brain

This is where the magic (and the pain) happens. We need to:

  1. Launch headless browser instances.
  2. Navigate to LLM UIs (e.g., chat.openai.com, claude.ai).
  3. Log in (manual first, then maintain session).
  4. Interact with chat elements (input box, send button).
  5. Scrape responses.
  6. Manage agent turns.

Key Tools:

  • express: For the HTTP endpoint from Flutter.
  • puppeteer-extra: Because puppeteer-extra-plugin-stealth is essential for not getting immediately detected as a bot.
  • ws: For WebSockets, if you want real-time streaming.

Here’s a simplified agentManager.js concept:

// backend/agentManager.js
const puppeteer = require('puppeteer-extra');
const StealthPlugin = require('puppeteer-extra-plugin-stealth');
puppeteer.use(StealthPlugin());

const agents = {}; // Stores Puppeteer page instances and agent states

async function launchAgent(agentName, loginUrl, chatInputSelector, chatOutputSelector) {
    if (agents[agentName]) {
        console.log(`${agentName} agent already running.`);
        return;
    }

    console.log(`Launching ${agentName} agent...`);
    const browser = await puppeteer.launch({
        headless: true, // Set to false for debugging UI interactions
        args: [
            '--no-sandbox',
            '--disable-setuid-sandbox',
            '--disable-gpu',
            '--disable-dev-shm-usage',
            '--remote-debugging-port=9222', // Useful for debugging headless browsers
            '--window-size=1920,1080'
        ],
    });
    const page = await browser.newPage();
    await page.setViewport({ width: 1920, height: 1080 });

    try {
        await page.goto(loginUrl, { waitUntil: 'networkidle0', timeout: 60000 });

        // TODO: Implement actual login logic here.
        // This is highly specific to each LLM's login flow.
        // For simplicity, assume you manually logged in once and saved cookies,
        // or that the page just redirects to chat after initial visit.
        // For real use, you'd need to handle CAPTCHAs, 2FA etc.
        console.log(`Navigated to ${agentName} at ${loginUrl}. Assuming logged in.`);

        agents[agentName] = { browser, page, chatInputSelector, chatOutputSelector, busy: false };
        console.log(`${agentName} agent ready.`);
    } catch (error) {
        console.error(`Error launching ${agentName}:`, error);
        await browser.close();
        throw error;
    }
}

async function sendMessage(agentName, message) {
    const agent = agents[agentName];
    if (!agent) {
        throw new Error(`Agent ${agentName} not found.`);
    }
    if (agent.busy) {
        throw new Error(`Agent ${agentName} is currently busy.`);
    }

    agent.busy = true;
    const { page, chatInputSelector, chatOutputSelector } = agent;

    try {
        await page.waitForSelector(chatInputSelector, { timeout: 30000 });
        await page.type(chatInputSelector, message);
        await page.keyboard.press('Enter');

        // This is the tricky part: waiting for the LLM's response.
        // Needs careful selector watching for new messages, or status indicators.
        // Example: Wait for a specific new div to appear or for 'typing...' to disappear.
        // This 'waitForResponse' would be a custom function.
        const response = await waitForResponse(page, chatOutputSelector); // Placeholder
        agent.busy = false;
        return response;

    } catch (error) {
        agent.busy = false;
        console.error(`Error communicating with ${agentName}:`, error);
        throw error;
    }
}

// Placeholder for a robust response waiting function
async function waitForResponse(page, chatOutputSelector) {
    // This is highly dependent on the LLM's UI structure.
    // You might need to watch for new message elements, or changes in existing ones.
    // Example: Watch for the last message element to change or a new one to appear.
    // For ChatGPT, you might look for the last assistant message.
    await page.waitForSelector(chatOutputSelector, { timeout: 60000 }); // Wait for the chat output div
    await page.waitForFunction(selector => {
        const el = document.querySelector(selector);
        // Look for common signs of a completed response
        // e.g., no '...' or 'typing' indicator, or a new message block fully rendered.
        return el && el.textContent.trim().length > 0 && !el.textContent.includes('...');
    }, { timeout: 60000 }, chatOutputSelector);

    const lastMessageHandle = (await page.$$(chatOutputSelector)).pop(); // Get the last message element
    if (lastMessageHandle) {
        return await lastMessageHandle.evaluate(el => el.textContent);
    }
    return 'No response found.';
}

module.exports = { launchAgent, sendMessage };
Enter fullscreen mode Exit fullscreen mode

And your app.js (Express server):

// backend/app.js
const express = require('express');
const cors = require('cors');
const { launchAgent, sendMessage } = require('./agentManager');

const app = express();
const port = 3000;

app.use(cors());
app.use(express.json());

// Initialize agents on startup
(async () => {
    try {
        await launchAgent('ChatGPT', 'https://chat.openai.com/', 'textarea#prompt-textarea', 'div.markdown.prose');
        await launchAgent('Claude', 'https://claude.ai/', 'div.ProseMirror', 'div.MessageComponent_markdownContainer__hQj_Y');
        // You'd add Grok, Gemini, etc. here with their specific selectors
    } catch (error) {
        console.error('Failed to launch agents:', error);
        process.exit(1);
    }
})();

// Simple endpoint for Flutter to interact
app.post('/chat', async (req, res) => {
    const { prompt } = req.body;
    if (!prompt) {
        return res.status(400).send({ error: 'Prompt is required.' });
    }

    try {
        // Simple round-robin or sequential agent interaction for demonstration
        const agentNames = Object.keys(require('./agentManager').agents);
        const firstAgent = agentNames[0]; // Start with ChatGPT for example

        console.log(`User prompt: "${prompt}"`);
        const response1 = await sendMessage(firstAgent, prompt);
        console.log(`${firstAgent} response: "${response1.substring(0, 50)}..."`);

        // Now, pass Agent1's response to Agent2
        const secondAgent = agentNames[1]; // Claude for example
        const response2 = await sendMessage(secondAgent, `Here's what ${firstAgent} said: "${response1}". What are your thoughts?`);
        console.log(`${secondAgent} response: "${response2.substring(0, 50)}..."`);

        // Send a combined or final response back to Flutter
        res.json({
            agent: 'Orchestrator',
            response: `User prompt initiated a chain: \n\n${firstAgent}: ${response1}\n\n${secondAgent}: ${response2}`
        });

    } catch (error) {
        console.error('Chat orchestration error:', error);
        res.status(500).send({ error: 'Failed to process chat with agents.' });
    }
});

app.listen(port, () => {
    console.log(`Node.js orchestrator listening at http://localhost:${port}`);
});
Enter fullscreen mode Exit fullscreen mode

Important: The selectors (textarea#prompt-textarea, div.markdown.prose, div.ProseMirror, div.MessageComponent_markdownContainer__hQj_Y) are examples and are highly prone to breaking when LLM providers update their UIs. This is a constant battle. The waitForResponse function is also a simplification. Real-world implementation requires much more robust logic to determine when an LLM has finished responding.

What I Got Wrong First

I hit a wall constantly with Puppeteer timeouts, especially when trying to automate ChatGPT v4. The UI sometimes loads in stages, or dynamic content shifts around elements. My initial await page.waitForSelector('#chat-input', { timeout: 30000 }); would frequently throw a TimeoutError: Waiting for selector '#chat-input' failed: Waiting for selector failed: timeout 30000ms exceeded even when the element visually appeared present.

Turns out, Puppeteer 21.6.1 had this specific behavior where waitForSelector could fail if the page was still heavily loading background assets or if the DOM was undergoing rapid mutations, even if the element was technically there in the snapshot. It wasn't truly ready for interaction.

The Fix: I had to switch to a more resilient approach using page.waitForFunction combined with network idle checks. Instead of just waiting for the selector, I waited for a condition involving the selector, ensuring it wasn't just present but also interactive or stable.

// Old, flaky approach:
// await page.waitForSelector('textarea#prompt-textarea', { timeout: 30000 });

// New, more robust approach:
await page.waitForFunction((selector) => {
    const el = document.querySelector(selector);
    return el && !el.disabled && el.offsetParent !== null; // Check if element exists, is not disabled, and is visible
}, { timeout: 45000 }, 'textarea#prompt-textarea'); // Increased timeout, more specific check

// And crucially, after interacting, wait for the network to be idle to ensure responses are fully loaded
await page.waitForNavigation({ waitUntil: 'networkidle0', timeout: 60000 })
    .catch(e => console.log('Navigation idle timeout (expected for some LLMs):', e.message)); // Catch expected timeouts
Enter fullscreen mode Exit fullscreen mode

This change alone drastically reduced the TimeoutError occurrences and made the nodejs cli automation much more stable.

Optimizing LLM Agent Cost Reduction

The llm agent cost reduction with this method is undeniable. But it's not without its own set of challenges and optimizations.

  1. Headless vs. Headed: Always use headless: true in puppeteer.launch for no api key ai agents in production. It consumes less memory and CPU. Only run headless: false for debugging UI interactions.
  2. Resource Management: Each browser instance consumes RAM and CPU. For a truly scalable flutter multi agent system, you'll need a beefy server or distributed architecture (e.g., using Kubernetes to manage browser pods). Monitor memory aggressively. Chromium can be a hog.
  3. Bot Detection: LLM providers are getting smarter.
    • puppeteer-extra-plugin-stealth is your first line of defense.
    • IP Rotation: If you're running many concurrent sessions from the same IP, you might get blocked. Consider proxy providers (e.g., Bright Data, Oxylabs) for rotating IPs. This adds cost but is cheaper than API keys.
    • User-Agent Strings: Rotate them. Make them look like real browsers.
    • Human-like Delays: Don't type instantly. Add await page.waitForTimeout(Math.random() * 2000 + 500); between actions.
  4. UI Resilience: LLM web UIs change. Frequently. This is the biggest maintenance headache. Your selectors (chatInputSelector, chatOutputSelector) will break. You need robust error handling and probably a way to dynamically update selectors or visually identify elements if a hardcoded selector fails. Honestly, it's a pain. But the cost savings make it worth the grunt work for specific use cases.
  5. Concurrent Sessions: A single Node.js process can manage multiple browser based llm chat sessions, but don't overdo it. Start with 2-3 per browser instance, then scale up.

This approach means you're not paying per token, but for the compute resources needed to run headless browsers. For high-volume internal tools or specific multi llm chatroom cli agents where the conversation length is unpredictable, this is a clear winner for llm agent cost reduction.

FAQs

Is this approach truly API-key free?

Yes, for the LLM interaction itself, you bypass direct API keys entirely. You're simply automating web browsers, just like a human would. However, you'll still need API keys for other services if your application requires them (e.g., Firebase, Stripe, other traditional APIs).

What are the major downsides of CLI-powered LLM agents?

The primary downsides are fragility and resource consumption. Web UIs are constantly changing, breaking your selectors and automation scripts. Maintaining these scripts requires continuous effort. Additionally, running multiple headless browsers consumes significant CPU and RAM, which can limit scalability and add server hosting costs.

How scalable is a Flutter multi-agent system using browser automation?

Scalability depends heavily on your backend's resource management and bot detection countermeasures. A single Node.js server can handle a limited number of concurrent browser sessions. For higher scale, you'd need to distribute your Node.js processes across multiple servers, potentially using a container orchestration system like Kubernetes, and employ IP rotation to avoid rate limits or IP blocks from LLM providers.

This isn't a silver bullet. You're trading per-token API costs for operational complexity and server overhead. But for specific multi llm chatroom cli agents use cases, especially where llm agent cost reduction is paramount and you can tolerate some maintenance, this no api key ai agents strategy is incredibly powerful. The 90% cost savings are real, not just marketing fluff. If you're serious about building something like this, hit me up on buildzn.com. I've been there, done that, and got the battle scars.

Top comments (0)