<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Prasad Chalasani</title>
    <description>The latest articles on DEV Community by Prasad Chalasani (@pchalasani).</description>
    <link>https://dev.to/pchalasani</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1150257%2F80abd9db-6756-4981-85b0-30c1fef7971d.jpg</url>
      <title>DEV Community: Prasad Chalasani</title>
      <link>https://dev.to/pchalasani</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/pchalasani"/>
    <language>en</language>
    <item>
      <title>Get Started with Local Language Models</title>
      <dc:creator>Prasad Chalasani</dc:creator>
      <pubDate>Sun, 17 Sep 2023 14:12:17 +0000</pubDate>
      <link>https://dev.to/pchalasani/get-started-with-local-language-models-e2l</link>
      <guid>https://dev.to/pchalasani/get-started-with-local-language-models-e2l</guid>
      <description>&lt;p&gt;I recently dove into the world of Local LLMs, to see how I can hook them up with &lt;a href="https://github.com/langroid/langroid"&gt;Langroid&lt;/a&gt;, the agent-oriented open-source LLM framework. These are some notes from what I learned, plus a pointer to a tutorial on how to use them with Langroid.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why local models?
&lt;/h2&gt;

&lt;p&gt;There are commercial, remotely served models that currently appear to beat all open/local&lt;br&gt;
models. So why care about local models? Local models are exciting for a number of reasons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;cost&lt;/strong&gt;: other than compute/electricity, there is no cost to use them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;privacy&lt;/strong&gt;: no concerns about sending your data to a remote server.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;latency&lt;/strong&gt;: no network latency due to remote API calls, so faster response times, provided you can get fast enough inference.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;uncensored&lt;/strong&gt;: some local models are not censored to avoid sensitive topics.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;fine-tunable&lt;/strong&gt;: you can fine-tune them on private/recent data, which current commercial models don't have access to.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;sheer thrill&lt;/strong&gt;: having a model running on your machine with no internet connection,
and being able to have an intelligent conversation with it -- there is something almost magical about it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The main appeal with local models is that with sufficiently careful prompting,&lt;br&gt;
they may behave sufficiently well to be useful for specific tasks/domains,&lt;br&gt;
and bring all of the above benefits. Some ideas on how you might use local LLMs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;In a mult-agent system, you could have some agents use local models for narrow 
tasks with a lower bar for accuracy (and fix responses with multiple tries).&lt;/li&gt;
&lt;li&gt;You could run many instances of the same or different models and combine their responses.&lt;/li&gt;
&lt;li&gt;Local LLMs can act as a privacy layer, to identify and handle sensitive data before passing to remote LLMs.&lt;/li&gt;
&lt;li&gt;Some local LLMs have intriguing features, for example llama.cpp lets you 
constrain its output using grammars.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Running LLMs locally
&lt;/h2&gt;

&lt;p&gt;There are several ways to use LLMs locally. See the &lt;a href="https://www.reddit.com/r/LocalLLaMA/comments/11o6o3f/how_to_install_llama_8bit_and_4bit/"&gt;&lt;code&gt;r/LocalLLaMA&lt;/code&gt;&lt;/a&gt; subreddit for&lt;br&gt;
a wealth of information. There are open source libraries that offer front-ends&lt;br&gt;
to run local models, for example &lt;a href="https://github.com/oobabooga/text-generation-webui"&gt;&lt;code&gt;oobabooga/text-generation-webui&lt;/code&gt;&lt;/a&gt;&lt;br&gt;
(or "ooba-TGW" for short) but the focus in this tutorial is on spinning up a&lt;br&gt;
server that mimics an OpenAI-like API, so that any Langroid code that works with&lt;br&gt;
the OpenAI API (for say GPT3.5 or GPT4) will work with a local model,&lt;br&gt;
with just a simple change: set &lt;code&gt;openai.api_base&lt;/code&gt; to the URL where the local API&lt;br&gt;
server is listening, typically &lt;code&gt;http://localhost:8000/v1&lt;/code&gt;. It really is as simple as that!&lt;/p&gt;

&lt;p&gt;There are two libraries I'd recommend for setting up local models with OpenAI-like APIs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/oobabooga/text-generation-webui"&gt;ooba-TGW&lt;/a&gt; mentioned above, for a variety of models, including llama2 models.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/abetlen/llama-cpp-python"&gt;llama-cpp-python&lt;/a&gt; (LCP for short), specifically for llama2 models.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you have other recommends, feel free to add them in the comments.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building Applications with Local LLMs
&lt;/h2&gt;

&lt;p&gt;We open sourced Langroid to simplify building LLM-powered applications, whether with local or commercial LLMs. If you’re itching to play with local LLMs in simple python scripts, head over to our &lt;a href="https://langroid.github.io/langroid/blog/2023/09/14/using-langroid-with-local-llms/"&gt;tutorial&lt;/a&gt; on using Langroid with local LLMs.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>llama</category>
      <category>tutorial</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
