DEV Community

Roshan Horo
Roshan Horo

Posted on Edited on

How ChatGPT Understands your Questions?

#ai

To Understand what ChatGPT does behind the scene, first we break the word “ChatGPT” into literal meaning - because OpenAI ( company who creates ChatGPT ) does the good job naming this product - as straightforward as it would be.

So, “Chat” means - we can interact with our natural language in the form of text or voice, if you prefer talking rather than texting.”GPT” - it’s stands for Generative Pretrained Transformer, sounds crazy right but in other words it’s a transformer which can generate or predict the next text based on the pretrained text ( lots of text ), will get into each technical terms in a bit and also It’s an LLM stands for Large Language Model build by OpenAI just like Claude,Llama,Gemini,etc.

But, how LLM can understand the text, it’a a machine right, it can only be understood by machine language that is binary (0s and 1s). How it can find the meaning into it and give the response that are close to it.

That’s where Transformer model comes into picture, so, we have large text broken down into small pieces so that it can plot into 3D space and like wise token are placed near to it. Like “places” token - “Paris” or “US” are close together.

The process of breaking down text into small pieces are known as Tokenization and ploting into 3D space to get meaning out of it is known as Embeddings and we use vector calculation for that.

Most of the Tokenization done for the same text are different based on LLM or models. Although, for some character like “,” or “.” they can be common (shown above).

So, when we put those token into an neural network to get the predicted token that would match the right token as close as possible, this model is known as Transformer model. It uses various mechanism like “Positional Encoding”, “Multi Head Self-Attention”, “Softmax” to get the right output.

How Tranformer works :

STEP 1 - Encoding : Divide large Text into smaller text and convert into numbers called tokens using llm vocabulary ( for tokens ) and plot that numbers into 3D map known as Embeddings.

STEP 2 - Positional Encoding : It's a mathematical process, so that we can get different meaning for the same sentences which has similar token. eg: "Tab is better than space" and "Space is better than tab", the meaning is changed based on position right.so, this is figured it out in this step.

STEP 3 - Multi Head Attension : To understand this,we need to clarify these terms -

  1. In first RNN ( Neural Network ) method - tokens like - "the River bank" and " the ICICI bank" is not able to grab the context what text is all about.

  2. In the Self Attension method - In this process, token are allowed to talk to each other and update their vector embeddings to get the correct context. This is also called Single Head Self Attension.

    Now, Multi Head Attension improves the contexual understanding by applying parallelization or by using multiple Head.Feed Forward means do the process again and again till the output is most predictable text.

STEP 4 - Linear & Softmax : Linear gives the probability distribution of the next token, by default it will the the high probability. Softmax is a function that you can choose linear of the less probability and get some different outcomes.

Top comments (0)