top of page

Welcome
to NumpyNinja Blogs

NumpyNinja: Blogs. Demystifying Tech,

One Blog at a Time.
Millions of views. 

LLM - Large Language Models

Feb 11
6 min read

Large Language Models (LLMs) are a type of machine learning model designed to understand and generate human language using deep neural networks. At their core, language models work by learning patterns in text and calculating the probability of the next word based on the words that come before it.


For example, given the input sentence “The sky is …”, the model predicts the most likely next word—such as “blue”—based on what it has learned from vast amounts of text data. By repeatedly predicting the next word, LLMs are able to generate coherent sentences, answer questions, summarize content, and engage in natural-language conversations.


Language Models


Category

Description


Large-Scale

Emergence

Era

Bag-of-Words Models

Represent text as a collection of unordered words, ignoring grammar, word order, and contextual relationships.

No

No

1950s-1960s

N Gram Model

Consider groups of N consecutive words to capture local word order and sequence information.

No

No

1950s-1960s

Hidden Markov Models (HMMs)

Represent language as a sequence of hidden states that generate observable outputs, capturing sequential dependencies probabilistically.

No

No

1980s-1990s

Recurrent Neural Networks (RNNs)

Process sequential data by maintaining an internal state, allowing the model to capture context from previous inputs.

No

No

1990s-2000s

Long Short-Term Memory (LSTM) Networks

An extension of RNNs designed to capture long-term dependencies in sequential data using gated mechanisms.

No

No

2010s

Tranformers

A neural network architecture that processes variable-length sequences using a self-attention mechanism, enabling efficient parallelization and long-range dependency modeling

Yes

Yes

2017-Present


Transformers: The Backbone of Modern Language Models


Most language models are built using the Transformer architecture. It is currently the best and most powerful way to teach computers how to understand and generate human language.


Popular tools like ChatGPT and Gemini are all based on Transformers. Every version of ChatGPT uses this same architecture. While the models keep improving, the basic Transformer design stays the same.


Transformers work well because they can look at all the words in a sentence at once and understand how they are related. This helps the model understand context better and give more accurate answers.


In short, Transformers are the backbone of modern language models and are still the top choice today.


Why Some Models Perform Better Than Others


Not all models perform equally on every task. The performance depends heavily on the data used for training:


  • Clean, high-quality data → Better performance

  • Noisy or low-quality data → Poor performance


For example, ChatGPT was trained on up to 1 trillion tokens, equivalent to around 10 million books—far more than any human could read in a lifetime (roughly 700 books).


What is a Large Language Model (LLM)?


A Large Language Model (LLM) is essentially a smart text engine trained on massive amounts of text data—articles, websites, and books. It can understand human language and predict the next word or phrase in a sentence based on context. Think of an LLM as a “brain” for text: it learns patterns and relationships in language to generate coherent responses.


The power of LLMs comes from neural networks and parameters:

  • Neural networks predict the next word in a sequence.

  • Parameters are the internal values that the model adjusts during training. The more parameters a model has, the better it can learn and provide accurate responses. For example, DeepSeek-71 has 7 billion parameters, allowing it to handle more complex language tasks.


As the number of parameters increases, the model becomes capable of serving more users and producing higher-quality outputs.


Tools for Running LLMs


Several tools and platforms make it easier to run, manage, and deploy LLMs:


  1. Ollama: Ollama allows you to run LLMs locally on your computer. It provides a runtime and model manager so you can download and use models like LLaMA2 or DeepSeek easily. You can run these models through a simple command-line interface (CLI).Website: ollama.com


  2. Hugging Face: Hugging Face is a platform and ecosystem for AI/ML models. Think of it as “GitHub for AI.” It provides:

    • Model Hub → Hosts over 500,000 models

    • Dataset Hub → Public datasets for machine learning

    • Spaces → Deploy AI apps using Gradio or Streamlit

    • Transformers library → Python SDK to use models easilyWebsite: huggingface.co


Understanding Tokens


A token is a chunk of text that an LLM reads or generates. It is not always a full word. Tokens can include:


  • A whole word (e.g., “cat”)

  • A subword (e.g., “play” + “ing”)

  • Individual characters (e.g., “a”, “b”, “c”)

  • Spaces or punctuation (e.g., “ ”, “,”, “.”)


Tokens are how text is broken down into pieces that models can understand. Using a token calculator, you can calculate the number of input/output tokens, words, characters, and total characters for your data.


Gemini Model Variants


Gemini offers multiple model versions: Gemini Pro, Gemini Flash, and Gemini Flash Lite. The underlying algorithms are the same across these models, but pricing and performance differ depending on the number of parameters.


LLM Agents


An LLM agent is like a robot. Unlike a standard LLM that only answers questions, an agent can plan, decide what to do next, and use other tools to achieve a goal or complete a task. Essentially, agents add reasoning and task-solving capabilities to language models, making them more interactive and powerful.


Types of Agents


Agents can be broadly classified based on their memory and decision-making capabilities:


  1. Simple Reflex Agents – No memory.

  2. Model-Based Reflex Agents – Memory of previous states.

  3. Goal-Based Agents – Act to achieve specific goals.

  4. Utility-Based Agents – Make decisions based on a utility function.

  5. Learning Agents – Improve performance over time through learning.


1. Simple Reflex Agents


Simple reflex agents act solely based on the current input and predefined rules. They do not remember past states or consider future consequences—they respond immediately to what they sense.


How They Work:

  1. Take an input, such as sensor data.

  2. Provide an immediate output based on predefined rules.

  3. Do not analyze past information or predict future outcomes.


Examples:

  • Room heaters that turn on or off depending on the current temperature.

  • Traffic light sensors that change signals based on real-time traffic flow.


Simple reflex agents are fast and efficient for straightforward tasks but are limited in handling complex or unpredictable environments.


2. Model-Based Reflex Agents


Model-based reflex agents maintain an internal representation of the world to make smarter decisions. Unlike simple reflex agents, they don’t just react to the current input—they also consider previous knowledge and the instructions they’ve been given.


How They Work:

  1. Take the current input from sensors or the environment.

  2. Refer to predefined instructions or rules.

  3. Use previous knowledge or past states to understand the world better.

  4. Analyze all this information to predict outcomes and decide the best action.


Example:A cleaning robot tasked with multiple rooms doesn’t just clean whatever it senses first. Instead, it plans: clean Room 1, then Room 2, and so on, efficiently completing the task based on its internal model.


By keeping an internal model, these agents can handle more complex and dynamic environments than simple reflex agents.


3. Goal-Based Agents


Goal-based agents make decisions with a specific objective in mind. Instead of just reacting to the current situation, they evaluate possible actions and choose those that bring them closer to achieving their goal.


Example:A self-driving car acts as a goal-based agent when its objective is to reach the destination safely. It considers various routes, traffic conditions, and potential obstacles, then selects the actions that help it achieve this goal efficiently.


Goal-based agents are more flexible and intelligent than simple reflex or model-based agents because they plan their actions based on desired outcomes.


  1. Utility Based Agent


A utility-based agent evaluates multiple possible outcomes and selects the one with the highest utility—that is, the outcome that offers the best overall result. Unlike simple agents that may act on fixed rules, utility-based agents weigh options and make decisions based on a measure of “how good” each possible action is.


Example: A self-driving car acts as a utility-based agent when choosing a route. It doesn’t just pick the shortest path; it analyzes various routes and selects the one that balances safety, speed, and efficiency to achieve the optimal outcome.


In essence, a utility-based agent doesn’t just act—it chooses the best possible action after carefully analyzing the available alternatives.


  1. Learning Agent


Learning agents improve their performance over time based on feedback from their actions and the environment. Unlike other agents that follow fixed rules, learning agents adapt and refine their behavior to achieve better results.


How They Work:

  1. Take an input or perform an action.

  2. Receive feedback on the outcome.

  3. Adjust future actions based on the feedback to improve performance.


Example:ChatGPT is a learning agent. When you provide a prompt, it generates responses and uses feedback—such as user ratings or corrections—to fine-tune its future answers, becoming more accurate and helpful over time.


Learning agents are especially powerful in dynamic environments where rules alone are not enough to achieve optimal results.



 
 

+1 (302) 200-8320

NumPy_Ninja_Logo (1).png

Numpy Ninja Inc. 8 The Grn Ste A Dover, DE 19901

© Copyright 2025 by Numpy Ninja Inc.

  • Twitter
  • LinkedIn
bottom of page