DEV Community

Kathir
Kathir

Posted on

DAY 1 - GENERATIVE AI

Generative AI

Generative AI refers to AI systems that learn patterns from existing data and use those learned pattern to generate new content such as text, video, audio, code and images.

Gen AI is a type of AI focused on generating new content.

Large language models are a major example of generative AI for text and code.

Generative AI is a field of AI that focuses on generating new content. Modern Generative AI system are primarily built using deep learning techniques, particularly architectures such as Transformers and diffusion models (used for image generation).

ML- Learns patterns to make prediction/decision.

DL- Uses neural networks with many layer to learn complex patterns.

GenAI - Uses learned patterns to generate new content.

LLM

  • LLM doesn’t look up the answer like Google.

💡

Prompt —> Tokenization —> Tokens —> Neural Network —> Predict next token —> Predict next token —> Predict next token —> … —> Final response.

  • An LLM, or Large Language Model, is a neural-network model trained on large amounts of text to understand and generate language. Given a sequence of tokens as context, it predicts subsequent tokens to produce a response.

API request format:

{
    "model": "some-model",
    "message": [
        {
            "role": "user",
            "content" : "Explain binary search"
        },
        {
            "role": "system",
            "content": "You are a Java programming tutor"
        },
        {
            "role": "assitant",
            "content": "Binary serach repeatedy divides a sorted array in half."
    }           
    {
        "role": "user",
        "content": "Give me Java code"
       }
    ]
}
Enter fullscreen mode Exit fullscreen mode
  • message ⇒ When you call a chat-based LLM, you give it a conversation history. Role and content are the conversation. The model reads this conversation and generates the next response.
  • There are three important roles:
    • system → defines how the assistant should behave.
    • user → Contains the user request.
    • assistant —> Represents previous response from the AI.

Reason for conversation history?

LLMs generally don’t automatically “remember” previous API calls.

A chat completion request contains a sequence of messages. Each message has a role and content. The role identifies the source or purpose of the message, such as system, user, or assistant. The content contains the actual text or other supported input. The model uses these messages as context to genrate the next assistant response.

💡

The quality and structure of the input can significanly influence the output.

Temperature

  • Low Temperature —> More predictable, More consistent
  • High Temperature —> More varied, More Creative.
  1. Should you use high temperature for a banking application?

    Generally, I would prefer lower randomness for task requiring consistent and predictable outputs, such as financial or structured application.

    For creative tasks, a higher temperature can be useful.

Context Window

The context Window is the maximum amount of tokenized information an LLM can process as context in a single request. It can include system information, conversation history, retrieved documents, and the current user query. In RAG, we retrieve only the relevant information and place it into the context so that model can answer without having to include the entire knowledge base.

Example: Chat with an LLM

Imagine you send information in a sequence of speaking like format. In the start you mention that you are coding in java. But after a set of conversation you are asking about the programming language coding now? So, something is required to remember that conversation about currently coding in java. So, that we required context window.

Hallucination

Hallucination is when generative AI model produces information that sounds plausible and confident but is factually incorrect, unsupported or completed fabricated. This happens because an LLM generates responses by predicting likely tokens based on patterns it learned, rather than inherently verifying whether every statement is true.

Hallucination can occur because the model is optimized to generate plausible responses rather than inherently verify whether every generated statement is factually correct.

How to reduce Hallucination?

  • Better prompting
  • RAG
  • Validation
  1. You are currently building an application that uses an LLM to generate Java code. Would you directly execute the generated code?

    No. LLM-generated code should be treated as untrusted input. First I would validate the code and check for security concerns, then execute it inside an isolated sandbox with strict resources and permission limits.

    Static validation is useful, but it isn’t sufficient as the primary security boundary because malicious or unexpected behavior can bypass simple checks. The actual execution should happen in a sandbox with least privilege and resource allocation.

    For example: I would use container isolation, disable unnecessary network and filesystem access, enforce CPU/memory/time limits and run security checks before execution. I would also capture the output and errors for validation.

    I haven't worked with Docker hands-on yet, but for executing LLM-generated code, I would use a containerized sandbox such as Docker to isolate the execution environment. I would restrict network access and impose CPU, memory, filesystem, and execution-time limits so that the generated code cannot affect the host system.

Potential concerns: malicious code, infinite loops, resource exhaustion, incorrect code, data access, arbitrary file access, network access.

💡

Internal Documents —> Chunking —> Embedding —> Vector Database

  1. You are building an AI chatbot for Amazon's internal employees. The chatbot must answer questions using internal company documents. The LLM's training data doesn't contain these private documents. How would you design the system?

    I would use a RAG architecture.

    Why can’t you just put all the company documents into the LLM’s context context window?

    Imagine, the company have 1 million document. Then, you can’t simply send all of them to the context window. The problems include: Context-window limitation, High token cost, Higher latency, Irrelevant information, Potential worse retrieval/answer quality, Privacy/access-control concerns.

    RAG says “Don’t give the LLM everything. First find the information relevant to this particular question, then give that information to the LLM as context”.

    You chose RAG. But how does your system know which document are relevant to the user’s question?

    I would convert the user’s question into an embedding vector and compare it with the embedding vector of document chunks stored in a vector database. Using similarity metrics such as cosine similarity, I would retrieve the most relevant chunks and provide them as context to the LLM.

Embedding

An embedding is a numerical vector representation of data such as text. It converts the semantic meaning of text into numbers so that similar pieces of information have similar vector representations. These vectors can then be stored in a vector database and used for semantic search or retrieval in systems like RAG.

Vector database —> Stores the vector generated by embedding model and helps to retrieve similar contents.

RAG —> Combines retrieval + LLM generation

Embedding model —> Converts text into vectors representing semantic information.

  1. How would you convert a 500-page PDF into something that can be stored in a vector database?

    500-page PDF —> Text Chunks —> Tokenization —> Tokens —> Embedding model —> Vector

    Or

    PDF —> Extract text —> Clean / Preprocess —> Chunk the text —> Embedding model —> Embedding vectors —> Vector database

    You have a 500-page document. How do you decide the size of each chunk? Would you simply make every chunk 500 tokens?

    We want chunks to be small enough for precise retrieval, but large enough to preserve the meaning/context.

    I wouldn’t simply divide the documents into equal parts. I would chose chunk size based on the model’s token limits and the nature of the document, while trying to preserve semantic boundaries such as paragraphs or sections. I may also use overlap between chunks so that important context isn’t lost at the boundaries. I would then evaluate different chunk sizes based on retrieval accuracy and latency.

    How does the system decide that Chunk 17 is more relevant than Chunk 42?

    We convert the user’s query into an embedding and compare it with the embedding of the document chunks using a similarity measure such as cosine similarity. The chunks with the highest similarity scores are considered more relevant and are retrieved as context for the LLM.

💡

The embedding represents the semantic information, and a similarity calculation between embeddings helps us measure relevance.

You retrieve the top 10 chunks. Do you send all 10 chunks to the LLM?

I wouldn’t blindly send all retrieved chunks to the LLM. First, I would check their relevance scores and select the most relevant chunks. Then I would make sure their total token count fits within the model’s context window, while leaving room for the system prompt, user’s question, and generated response. This also reduces latency and token cost.

  1. Suppose you test your RAG system with 100 questions.

You discover:

100 questions
↓
80 retrieve the correct document
20 retrieve the wrong document
Enter fullscreen mode Exit fullscreen mode

Then among those 80:

80 correct retrievals
↓
70 correct answers
10 wrong answers
Enter fullscreen mode Exit fullscreen mode

"Where is your biggest problem: retrieval or generation?"

It should be on retrieval. Since,

Retrieval accuracy = 80/100= 80%

Generation accuracy = 70/80 =87.5%

I would investigate retrieval first because only 80% of queries retrieve the correct information. Since, the LLM cannot reliably answer a question if we provide irrelevant context, improving retrieval should have the biggest impact. I would analyze the embedding model, chunking strategy, similarity metric, Top-K selection, and potentially introduce a reranking stage.

RAG —> Retrieval ⇒ Embedding Vector DB

—> Generation ⇒ LLM

  1. Why do we need an embedding model? Why can't we just send the user's question directly to the vector database?

    If we search directly using user’s text, we would typically rely on lexical or keyword-based matching, which may miss documents that express the same meaning using different words. An embedding model converts both the user’s query and documents into “numerical vector representation” that capture semantic meaning. We can then compare the query vector with document vectors using a similarity measure such as cosine similarity and retrieve the most relevant chunks.

Cosine Similarity

How similar is the direction of two vectors?

For two vector: A=[a1,a2,a3,..] and B=[b1,b2,b3,…]

Cosine similarity compares their direction.

Smaller angle = More similar, Larger angle = Less similar.

Hands-on

SentenceTransformer is a Python Library interface/class used to create text embeddings. It comes from the Sentence Transformers library. It is a library/class that lets you load different embedding models.

Sentence Transformers is a library for generating embedding, and it supports different pre-trained embedding models. The specific model determines the resulting vector representation.

1. from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim

2. documents = [
    "Amazon leaders start with the customer and work backwards.",
    "Amazon employees should take ownership of their projects.",
    "Binary search works efficiently on sorted arrays.",
    "Java uses classes and objects for object-oriented programming."
]

3. query = "Tell me about customer-first leadership"

4. model = SentenceTransformer("all-MiniLM-L6-v2")

5. document_embedding = model.encode(documents)
6. query_embedding = model.encode(query)

7. similarities = cos_sim(query_embedding, document_embedding)[0]
# Give me the similarity scores for the first query. So, that it is [0]

8. for document, score in zip(documents, similarities):
    print(f"{score:.4f}->{document}")

9. best_index = similarities.argmax()

10. print("\nMost relevant document:\n")

11. print(documents[best_index])

Enter fullscreen mode Exit fullscreen mode

Line 8: This is a Python f-String used to format and print similarity score along with the document. f means formatted string. It allows you to put variables directly inside {}.

{score:.4f} —> : denotes formatting begins and .4f display as a floating point number with 4 digits after the decimal.

Line 7: In this line we use [0] since it represents the first query similarity. If there is more than one query we could take [1] [2] like that.

[
[0.82, 0.31, 0.05], ← Query 1
[0.20, 0.75, 0.10] ← Query 2
] → Similarities

Line 9: argmax() means argument of maximum → it returns the position (index) of the largest value, not the largest value itself.

argsort(descending=True)[2] —> To find Top 2

top_index = similarities.argsort(descending=True)[:2]

print("\nTop 2 relevant document:\n")

for index in top_index:
    print(documents[index])
Enter fullscreen mode Exit fullscreen mode

Multiple queries:

from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim

documents = [
    "Amazon leaders start with the customer and work backwards.",
    "Amazon employees should take ownership of their projects.",
    "Binary search works efficiently on sorted arrays.",
    "Java uses classes and objects for object-oriented programming."
]

query1 = "Tell me about customer-first leadership?"
query2 = "What is the ownership principle?"

queries = [query1, query2]

model = SentenceTransformer("all-MiniLM-L6-v2")

document_embedding = model.encode(documents)
query_embedding = model.encode(queries)

similarities = cos_sim(query_embedding, document_embedding)

print("Query 1:")
for document, score in zip(documents, similarities[0]):
    print(f"{score:.4f}->{document}")

print("Query 2:")
for document, score in zip(documents, similarities[1]):
    print(f"{score:.4f}->{document}")

print(f"{query1}\n")
print(f"{documents[similarities[0].argmax()]}\n")

print(f"{query2}\n")
print(f"{documents[similarities[1].argmax()]}\n")

Enter fullscreen mode Exit fullscreen mode

Top comments (0)