Saturday, August 08, 2026

Expert Systems to Generative AI — tiny steps that caused giant leaps in productivity

In the beginning, we wrote the rules by hand. The expert systems of the 1980s and 90s were magnificent and exhausting. You found the best person in the building — the diagnostician, the underwriter, the network guru — and you sat a knowledge engineer across from them for months, extracting their judgment one IF-THEN at a time. Thousands of rules, curated into a knowledge base, executed by a deterministic inference engine that never hallucinated, never improvised, never surprised you. The execution was flawless. The problem was everything else: knowledge engineering cost a fortune, the expert could never quite articulate the intuition they actually used, and the systems were brittle in a way that bordered on comic — one step outside the rulebook and the magnificent machine went silent, or confidently applied a rule that did not fit. Every new product version meant another expensive rule-surgery project. The support costs compounded, and expert systems became a cautionary tale: perfect execution, starved by its own knowledge.

The generative AI era inverted exactly one thing: where the rules come from. Instead of interviewing the expert, you curate the raw record of their work — every decision, every correction, every recovery — and let an autoregressive probabilistic model learn the patterns from the data. The rules are no longer written; they are learned. And here is the part people get wrong about how these systems run in production: the model does not execute anything. It fills in a carefully formatted card — a JSON form, a tool call — and hands it to a deterministic harness that reads the card and does the work. It reminded me of nothing so much as the mainframe punchcard. Judgment happens in the model; execution happens in code; the card is the boundary. The expert system is back, but the knowledge base is learned, and the inference engine is split in two: a probabilistic form-filler in front of a deterministic executor.

Talk is cheap, so I ran the experiment on myself. I took a small open-weights model — free to download — and trained a LoRA adapter on three months of my own agent sessions: every failure, every recovery, curated into a few thousand clean examples. The adapter training cost me under $50 in GPU time. The model itself was free, so the total cost of my custom model is $50. For this price nowadays you can get dinner in a good restaurant. But I have a model and it works. I ran it in ollama, served it locally as an endpoint, and pointed my harnesses at it — Interpreter, VSCode+RooCode. I could not tell the difference.

Your mileage may vary. But the direction is set: the cost of a custom model is now the cost of a dinner, the data is the work you already did, and the harness does not care whose card it reads

Thursday, August 06, 2026

Which task actually needs AI?

AI earns its place at the workplace when a task is a workflow, not a single computable step: data collected from scattered places, analyzed, and turned into a conclusion someone can act on. That workflow exists in every industry — but I am only familiar with IT, so the table below shows how it applies in IT. The pattern I find: AI collects, analyzes, and drafts the conclusion; the human owns it.

Task / Activity Human value add Where AI earns / helps
Report generation from data spread across enterprise databases and files Knows which questions matter and what the numbers mean for the business Collects the scattered data (retrieval), analyzes it, drafts the report in the house format
Infra cost-saving initiatives (find waste, rightsize, decommission) Owns the risk call, approves destructive actions, adjudicates exceptions Finds waste candidates, sizes the savings, drafts the justification a manager reads
Architecture evaluation against best practices Judgment under tradeoffs; the final call and the accountability Checks designs against conventions, cites the violated principle, drafts the review
Incident triage & error recovery Novel failure modes, escalation judgment, the 3 a.m. call Recognizes known failure modes and proposes the known fix instantly (my $50 specialist)
Capacity planning Business context — launches, seasonality, risk appetite Crunches the utilization series, narrates trends and the outlook
Compliance & audit evidence Interprets gray areas, signs the attestation Gathers evidence across systems, drafts the control narratives

Read the columns as a split of labor: the middle column is judgment and accountability; the last column is collection, analysis, and drafting. That is the whole claim. The tasks where AI helps most are not the ones with the most data or the fanciest math — they are the ones where the conclusion must be composed, in a specific voice, from scattered inputs, repeatedly.

Two consequences:

1. The human column is not ceremonial. Every row ends in a person who owns the outcome — the architect's call, the manager's approval, the attestation signature. Not caution theater: the model drafts, the human decides. The day a row loses its human value-add, it stops being an AI task and becomes a cron job — and that is fine too. Automation is the final stage of a well-understood task, not a failure of AI.

2. The AI column is smaller and cheaper than advertised. Judgment + your format + repetition is exactly the profile where a small model trained on your own data beats a frontier subscription. You do not need the model that writes poetry. You need the one that knows your loop.

Saturday, July 25, 2026

Kimi-K3 shift focus to Memory

AI inference economics is getting even more difficult. Newer models are compute efficient but memory heavy. This is shifting the economics in favor of hyperscalers and away from DIY labs like I have. There is no way you can build a lab in your garage to serve kimi-k3 (after they release their openweights next week). Forget ollama pull kimi-k3... 

If you are a financier of hyperscaler (buying their debt), this is good news because DIY and bottom feeders just discoved a BTE (barrier to entry). But if you are a lab want to further research, you are SOL'ed. 

Kimi-k3 has 2.8T parameters that all need to be in memory during inferencing. That memory requirement is approximately 1.7TB of DRAM. 

It's not just the weights — context eats memory too

That 1.7TB is just the model weights sitting in memory. There is a second memory cost that scales with your traffic: the context window. Every conversation with the model lives in memory as the KV cache — key and value vectors computed for every token, at every layer, that must stay resident in memory for the entire session so the model can attend back to them.

This means your memory bill grows with (a) context length and (b) concurrent users. One long-context session can chew through tens of GB of KV cache. Multiply by thousands of simultaneous users and the KV cache can rival the weights themselves.

So the DIY math is even worse than the table suggests. You don't just need enough memory to hold the model — you need enough to hold every user's conversation at the same time. Hyperscalers spread that cost across a fleet; your garage lab eats it all at once.

Precision FormatMemory Size
Platform
MXFP4 (Kimi's own FP4)~1.5-2 TBMultigpu cluster (single rack at least)
Q4 Quantization~1.0-1.2 TBSingle Enterprise Server with multiple GPUs
Aggressive 2-bit~1TBMight fit in a high end laptop, but performance is abysmal and lots of hallucinations
BF16/BF32~6-12TBMultirack solution - Only option is hosted
The Sparsity Opportunity

  • Current consumer and even most enterprise GPUs have limited hardware support for sparse operations.
  • Sparse models often require custom kernels and careful integration with inference engines.
  • Quality can vary dramatically depending on how aggressively sparsity is applied.

A 8GPU host costs around 700K today, to host the high precision model, we will need 3-5 racks of these hosts (assuming 2 per rack) to get 12-16TB in the cluster. The cost of that is approximately $6M. Add to that cost of power and cooling and we are talking $7-8M. Now imagine you have a great new model optimization, you will need to raise 50M+ just in seed round to get started to support a multi million burn rate. 

There is one technical path that could meaningfully change this equation: sparsity.

Sparsity techniques prune weights to zero during or after training, then use specialized kernels that skip calculations and storage for those zeros. In theory, a 2.8T parameter model with 70–90% sparsity could reduce its memory footprint by 3–5× while preserving most of its capability.

If the Kimi-k3 weights are released with high-quality sparsity patterns (or if the community rapidly develops them), the effective memory requirement could drop from 1.7 TB down to the 400–600 GB range. That would bring the model back into reach for well-equipped single-server labs instead of requiring an entire rack.

However, significant challenges remain:

The lab that cracks efficient, high-sparsity inference at this scale will enjoy a massive cost advantage. Until then, the economic headwinds strongly favor those who can amortize millions of dollars of infrastructure across vast workloads — the hyperscalers.



Saturday, May 30, 2026

Prompt is the new Config, Workflow is the new Product

 Prompt are dynamic configurations that can adapt to incoming requests. Today configurations are static files which may miss a lot of corner cases. Enter an LLM and now we can reason with APIs and dynamically workaround blocks. This one ability of the LLMs can create panic among all gateways to enterprise. A gateway allows/deny based on a static policy. What if we embed this intelligence in them so they adapt continuously to incoming traffic. This is atleast one reason why firewalls/cybersec products are receiving scrutiny from CIOs. A prompt injection can completely change the security posture of an enterprise. 

When you have a wild and powerful entity on the loose, you need to build harness to channel this energy. That is the workflow. More and more startups are now focusing on workflows instead of just functionality. They called this harness a wrapper around LLM. These wrappers are the new gateways into the enterprise. 

Tuesday, May 26, 2026

Token Pricing: Reasoning Tax,, GPU Utilization & GPU Recency

 According to newly published TPI (Token Price Index), average cost of 1M tokens is just over $2. It rose over 75% in 1 year. So does a model serving provider make money off this pricing? Yes and No. The marginal cost of power + infra + facilities ranges between 6 cents to 1.60 cents per million token. The key factors driving costs higher are (a) low utilization of GPU (currently around 15%), (b) reasoning tax (the token generated inside to support customer tokens (c) GPU recency. 

On the latest GPU, the cost of generating token is lower creating an incentive for providers to move up to the latest GPU. But this also means they pay for supply constraint which will likely stay forever because the GPUs that are one generation behind are not used. So the amortization on those GPUs which assume 100% consumption for the full 6 years is just too low i.e. margins are artificially high. 

The GPU utilization relies on many techniques at deployment time but batching is the one which makes a meaningful impact. Assuming 100% of the input is batched is euphoric. We are average 15-25% on recent GPU and almost 1-2% on older GPUs. No one wants to run on legacy GPUs. Net net this factor is also artificially inflating margins. 

The reasoning tax comes from ordinary queries but especially from agentic AI. Agents generate output that is sent back to the model which then starts generating reasoning tokens which are currently not charged to the customer. These reasoning token cost is eaten by the provider and is referred to as "Reasoning Tax". 

Saturday, April 11, 2026

OpenClaw is MicroSaaS

 This OpenClaw has a annoyingly long setup, but after all the effort, it is worth it. I set it up to routinely check my gmail. The process is well documented on openclaw but it does require account with Google Cloud, Installation of gcloud SDK, Go, GoG. And of course a lot of knowledge of how Linux works. I did it on WSL on windows. 

Finally, it gave me all my emails which is not the point. You can ask it questions which you can't in a email client. Check this out. 

Question: Who keeps sending me email? In the last week who has sent me the most emails? 



Question: Can you summarize all the emails from Federal Reserve Board?



Question: Can you summarize the April 8th FOMC Minutes? 


Ok, so all of this is pretty straightforward. But note that I haven't used any tokens or have I?  What about monetization? Who is making money with all this? 

It turns out LLM providers like Anthropic, Gemini are charging fees for tokens usage. Cloud hosting services charge for any managed instances (I am using my laptop). ClawHub skills for specific vertical (read custom integration) get paid. Skills vendors are making some money (not enough to quit your job yet but good for Happy Meals from McD). 



Wednesday, December 24, 2025

Local LLMs does not cut it

 After using ollama on a local LLMs, I found it is not very useful and I was constantly going to online versions of the same model or proprietary models. My local setup was under $1K and the results from the various queries over last 6 months were not any better than a web search. There are two problems with local inferencing (I am going to skip the most obvious one i.e. the hardware is not as powerful). 

The first one is distribution shift where the data on which the LLM is trained is not the same and sometimes differs significantly to the prompt query. This results in incomplete or superficial answers to queries where online models provide significantly more detail. This is kind of like talking to a person who is not in the know but knows enough to sound dangerous. 

The second is overfitting. This is where the weights have decided to converge on an answer for any set of prompts that are similar. It does not find accurate answers that fit the prompt so it serves the generalized version. This is kind of like thinking in stereotypes.

The key point is the real slim shady (yes eminem) here is the data itself. If you train on prior exams and then take an exam that was set by some external body, you will find yourself not as prepared to take the exam as you thought you were. The content and theories have not changed, just the way to test is different. Failure to detect and correct these shifts can be disastrous because sometime it means the model has to be trained from scratch. Yep that $150M used to train the model was flushed down the toilet. Hopefully, in 2026, we focus more on the data and how to partition and reassemble weights and not as much on GPU/VRAM. The latter is really a commodity and in abundant supply. 

BTW, B200 only supports certain versions of PyTorch. If this continues we will get fragmentation in jobs as with every new release of GPU, all software would have to be upgraded. In the end GenAI is not much different from Natural AI. 

Happy Holidays. 



Friday, May 23, 2025

Local LLM using Ollama and open-webUI

 I have a local server with Nvidia GPUs which I bought off ebay for $800. The GPU are RTX but there are 4 of them in the server. I run ollama on it and downloaded a few models that I use mainly to ask them how to configure other applications. I have multiple laptops where I run docker containers with open-webui and ollama-webui as shown below. Open-webui allows saving chats while ollam-webui does not. 



The performance is about as good as any online LLM server. I routinely get around 20+ Tokens/sec i.e. it is not a long wait to get your queries answered and stop and add information. In other words, it is usable. 




All the models are free and open source - so I pay nothing to use them. The only out-of-pocket expense was this server that I purchase. I could have made one with $99 2 socket intel motherboard (from Aliexpress), but I did not have time, so I bought it off some student off ebay. 

The answers are accurate and not all that different from using online version. Give it a try! It is fun and really really cheap. 


Appendix

I got the setup and asked it to write in a blog fashion all the steps necessary. I haven't validated them, but they look roughly what I did. 


 Title: Setting Up OLLAMA and Open-WebUI Across Server and Windows Laptop


In this blog post, we'll walk through the process of setting up OLLAMA, a powerful language model, on a server equipped with four NVIDIA GPUs, and connecting it to Open-WebUI, an intuitive user interface for large language models, on a Windows laptop using WSL (Windows Subsystem for Linux).


**Part 1: Server Setup**

First, let's set up OLLAMA on the server. Begin by updating the package lists and installing required dependencies:

```bash

sudo apt-get update && sudo apt-get upgrade -y

pip install ollama[all]

```

Now, we'll download and pull your favorite models from Hugging Face Model Hub. Here, we'll use DeepSea, Mistral, and QWEN:


```bash

ollama pull deepseek

ollama pull mistral

ollama pull qwen

```

To configure OLLAMA as a service, create a systemd unit file in `/etc/systemd/system`:

```bash

sudo nano ollama.service

```

Add the following content and save the file:

```ini

[Unit]

Description=OLLAMA Service

Requires=nvidia-smi.service

After=nvidia-smi.service


[Service]

User=<username>

ExecStart=/usr/local/bin/ollama start

Restart=always

EnvironmentFile=-/home/<username>/.ollama/config

WorkingDirectory=/home/<username>/.cache/ollama


[Install]

WantedBy=multi-user.target

```


Next, enable and start the service:


```bash

sudo systemctl daemon-reload

sudo systemctl enable ollama

sudo systemctl start ollama

```

Configure OLLAMA to use the GPU devices and expose it over a public IP using port forwarding. Update the `/home/<username>/.ollama/config` file accordingly:

```bash

# In [general] section

port = 8000


# In [gpus] section (add the necessary GPU IDs)

gpus = 0,1,2,3

```


**Part 2: Windows Laptop Setup**


Install WSL and Ubuntu if not already done. Open a Ubuntu terminal and update the package lists:


```bash

sudo apt-get update && sudo apt-get upgrade -y

pip install openwebui

```

Connect to the OLLAMA service on the server using the public IP and port:


```bash

openwebui connect <server_public_ip>:<port>

```


Once connected, Open-WebUI will launch, providing a user-friendly interface for interacting with your language models.


By following this guide, you've successfully set up OLLAMA on a server with multiple GPUs and connected it to Open-WebUI on a Windows laptop using WSL, enabling seamless access to powerful AI models from anywhere.


Saturday, February 01, 2025

DRS1 = DSV3 + GRPO + VR

 DSR1 is actually a reasoning model which also does chat. But the surprise that it outperformed o2 and others is because nobody paid much attention to this company and its publications. After the DSR1 announcement, I found these two papers: DeepSeekMath (April, 24) and DeepSeekCoder (June, 24). Had I read this last year, I would have been waiting for the actual DSR1. In fact, the real contribution was in V3 model from December, 24. DSR1 is easier to reproduce from DSV3. Getting to DSV3 is what is difficult and amazing that used older and crippled infra. (2.8M GPU hours) or roughly 2.8M*$3.5/hr = $9.8M. Compare that to several billions spent in pre-training current proprietary and open source models. 

To get to DSR1, you start with DSV3 and apply GRPO (defined in DeepSeekMath paper). It is not SFT with HF, it is RL with verified rewards (VR) - essentially no humans involved. The key contribution of GRPO actually came out in April in DSMath paper that is linked above. 

To reduce cost, they use FP8 and DualPipe algorithm which helps them reduce GPU memory consumption. Recall successive GPUs from both Nvidia and AMD are simply adding more HBMe memory to the GPU module. Also recall the H800 was crippled by more than halving its bandwidth to the VRAM (memory on GPU module). The training parameters they optimized on are compute to communication ratio and near zero all-to-all communication in a GPU cluster. They avoided tensor level parallelism and focused on traditional pipeline bubble removal. They developed middleware that optimized on inter-node communication that understood the underlying transport as IB or NVLink. 

In summary, all the cost efficiencies are standard HPC techniques and their key innovations were published months before they achieved DSR1. The only difference is they believed in their approach and those who actually read the papers in April, 2024 were high on we need more GPUs to pre-train and ignored their innovations. 

The constrained infra in China pushed them in this direction, but not we should thank that constraint because we now know that even models exhibit emergent behavior when trained in constrained environment. You want sweet grapes, don't water the vine!


Friday, January 31, 2025

DeepSeek is On-Prem and "" real value is between ""

 High Flyer is a hedge fund which was early adopter of GPU acceleration in finance. The expertise they built in that field helped them launch a subsidiary Deepseek which recently released R1 after two other LLMs. Everybody now knows what Deepseek is but not many know that it is actually not running in any cloud. It is all on-prem.

In China, you can't get H100, but you can get H800 which is BW limited and H20 which is crippled version but enough to train and infer 680B parameter model. Some suggest that Deepseek innovated due to constraints. It may be, but it could just as well be that they figured out first that one does not need to load the whole model into memory for inference. The latest figures are that DSR1 only load 30B parameters which you can load on a Laptop GPU with only 12GB of VRAM. 

Since the new on DSR1 hit the wires, the consumer grade GPUs have been flying off the shelf. It is super easy to run DSR1 using ollama (literally just type %ollama run deepseek-r1). Most of the value, I derive from DSR1 is the bit inside. For example, I asked DSR1 how to compress and load a LLM into VRAM. The answer was obvious as shown below:  

In summary, my approach would be:
1. Start by quantizing the model parameters to lower bit precision (e.g., 8-bit or 16-bit).
2. Apply pruning to remove unnecessary weights after quantization.
3. Use techniques like dynamic or static quantization for further compression during inference.
4. Implement checkpointing strategies to optimize memory usage during the inference process.
5. Possibly combine with other methods like knowledge distillation if there's excess capacity.

but the real insight was in between tags <think>. Some of the tips there are knowledge distillation (not the same as model distillation. Use of fixed point math (something we did like 30 years ago to draw pictures using postscript). I particularly liked the way it went from top-of-the-head response of quantization to fixed point to knowledge distillation and went on to (eventually reject) sparsity techniques. 


Sunday, October 27, 2024

AI IDEs - Do you need it?

 AI generated code IDEs like Replit, Cursor.sh and plugins into vscode are all the rage now a days. I blogged earlier on using continue.dev in vscode and using online model for code generation. 

They are all priced between $12-$20/month and require a subscription (API Key) to a inference engine (OpenAI) on top of that. You can avoid all these fees, if  you run your model locally. I tried that using ollama locally on my laptop and pick a 7B or smaller model (I like codestral). It is slower but completly usable. 

That works well for hobbyist but what about teams in software factories? Writing code as a group is different workflow. GenAI powered tools will need to evolve to fit the team workflow. e.g. doing a diff between two code bases. Increasing the code context to include the whole team and not just my code base. 

There is also a thing I noticed, it works well for 3rd and 4th generation languages, but not for assembly/C. It does not understand compiler optimizations yet. Charging $20/month for this early alpha type functionality is not worth it imho. You can get equivalent general purpose code generation for free without seeing this error 


The impressive part of code IDEs is not that they generate code, it is that they can summarize your code for comments and audits. That may be the use case we want to focus on. 

Saturday, September 14, 2024

AI AlterEgo

 The killer application for AI is to enable expert profiles in enterprise and productivity applications. These are not bots that help you get through mundane tasks, these are profiles that application consults to provide assistance in using the application. This is akin to expert levels in gaming. 

When using - say - an IDE. Today, the profile that the application stores on the user is mainly to collect credentials and secure access to outside storage and other artifacts. With GenAI, these profiles can be based on other users, experts or just AI itself. If you admire writing style of someone, then assuming the person is willing to sell/share their profile, one can use that profile (import it) and the application can now use the profile's style to generate your content. 

IDE is the easiest to understand this concept, but you can imagine how intelligent profiles can be used in every sphere where applications perform majority of the mundane tasks. For example, the world's best trader can export his/her profile in stock trading and you can use if in your brokerage application to receive recommendations for trade which otherwise you would not have entered into as you would not have seen the opportunity. 

All of this is essentially creating a alter ego of yourself. Now everybody can become a rock star!

Saturday, August 03, 2024

Where is the productivity in AI? Try this!

For some reason, mainstream is now asking for proof of productivity from AI. There are some skeptics. Let me show you how it increases my productivity as a developer. 

As Easy as 123

Using continue code assistant, I was able to build with very little help an application that uses streamlit for UX, MySQL for DBMS and LangChain for chaining model and logic. The ease with which I can now "talk" to my tables in DBMS makes mysql workbench kind of obsolete. For run-of-the-mill DBMS reporting, we don't need to use any expensive human talent to get it done. This is a boost in productivity for anyone who has to back up an argument with data. This is probably why Snowflake acquired streamlit. The productivity gain is astronomical as I don't have to keep searching for the "right" syntax. I barely check on API reference as the code tells me which method and object I need. 

Agents R US

While I used chains and linked them together to get the end result, I could have created distinct semi-autonomous agents which would get the information as it updates and report the state in real-time. This type of work takes weeks today in a organization and it can now be done in hours. 

PaaS This!

You need a source of data, a connector to read/write to the data and a execution environment which allows for use of models from many sources (using their API key) and frameworks to keep all these components and their state in sync. This is not done in a IaaS setting, this needs a PaaS. No wonder, you can't get away from HuggingFace. No wonder they just announce Github models as competition to HF. 

Models are 4Ever: 

APIs come and go, but models are forever. I have used three different models in a single application and am paying no more 2.5 cents for 10K tokens. I am beginning to wonder if they actually make money providing me this service. Let's look at their investment, a typical model (like GPT), requires 144GPUs to load GPT model, it uses 750W per GPU used. Most cannot afford this, so they go for a model that fits into a single system, but single system needs to be configured for GPU passthrough to the VMs without any bloatware from K8s. We are looking at a hard requirement that a single node offer performance of 200TFLOPs as a minimum. 

Larger models with 400B+ parameters are now called giants. We need Giants because they keep the context around for longer and capture deeper relationships between parameter. But these giants shouldn't share context across tenants. I believe currently they do. 


Saturday, April 20, 2024

Llama 3 - More ways to run it, but still nothing new

 Llama 3 is out and getting to it can be a challenge. The approval email's URL expires in 24 hours. It can take 8hrs to download. But after the download from Meta, it can be use locally in text-generation-webui. This time it has hosted versions on hugging chat and meta itself. It says it's training stopped in 2021 so it continues to think the PM of UK is Boris. But it believes it is more conversational. 





When asked how many params it is trained on, it initially said 1.5B. Then I asked again and it changed its mind. 



Using ollama to run llama-3, I get better answers



On text-generation-webui, the model does not load except when you pick transformers as the loader. And the chat is not fully functional. 




After converting to GGUF, 



LM Studio is the best one of these for now. 




Thursday, April 04, 2024

LLM - Not everything can be learned - so let's realign it to our preferences

When I first started researching LLMs it seemed like the technology could simply learn and get to a point where it is self-learning artificial lifeform (AGI). Now 8 months since my last post, it looks like the initial trust of teaching an LLM everything is not giving the returns that researchers thought. Words originally used such as "emergent behavior" are now being replaced with "hallucinations", "catatrophic degradation". 

The jack of all LLM is not what we really wanted, what we want is precise control over the completions (answers). To get there, we are now seeing new aveneues of research collectively called fine-tuning. Fine-tuning is not a performance run-time effort, rather, it is changing the model's weights to reflect preferences. A new alphabet soup of acronyms called DPO, IPO, KTO are all optimizations that introduce new labeled datasets and under supervision get a generic pre-trained model to answer the "money questions". 

If you have been exposed to ML/AI for long, you already know we have seen this before and then it was called "reinforcement learning". Today they add a HF (human feedback to it) and it is now called RLHF. Once again, we are back to using likelihoods (read probabilities) and rewards (biases) to get an AI to spit out answers which can add economic value. 



Saturday, July 22, 2023

Costs in Training LLMs

 I went through the Llama-2 white paper that was released with the model by meta. I was hoping to learn some special technique they may be using to train their models. Apparently, there isn't any. The learning process is straightforward. What is different is the huge costs associated with fine tuning after the model is trained. This fine tuning requires human interaction and feedback. To incorporate the feedback, the model has to be altered that requires more computation. Training, fine tuning the model costs more than $20M (~$4 per hour and 5M hours). This immediately limits the number of players who will actively develop LLMs. The cost of adding safety to these models (e.g. block prompts for ransom letters etc.) is almost as high as cost of training the model. 

Another interesting tidbit from the paper was the assertion that RoCE based 200 Gbps interconnected cluster was adequate and more economical than Inifiband based cluster. RoCE uses commodity ethernet. If one can train 70B parameter model with trillions of tokens using commodity ethernet with RDMA, what is the compelling need to move to expensive NVLink linked superchips based systems? May be they are overfitting? (pun intended)

There is a significant cost to building these models that are shared with public (unknowingly) i.e. they are climate related. The carbon emission of these clusters is shown in the paper at 539 tonnes of CO2e. It took 3.3M hours of GPU (A100-80G). All of this to chat with a bot?

I found more benchmarks and metrics related to safety, climate and other social concerns in the paper than what one finds a technical paper. 

It was easy to play with the model using oobaooga's text gen UI. I used the 13B parameter model from the family of Llama-2s released. It is a bit dated. You can see for yourself. 




Sunday, July 09, 2023

Learning a Model

 Neural Networks have a bad reputation of being very confident when they are wrong. This is the result of a bad probability estimates being calculated (i.e. learned). They also suffer from adversarial attacks. Training is the activity that takes the most time in arriving at a functional LLM. Besides collecting, curating and integrating data sets, we have to also navigate around pot holes by employing optimization techniques on the objective function. Objective function or goal seeking function is a function that takes data and model parameters as arguments and outputs a number. The goal is to find values for these parameters which either maximize or minimize this number. Maximum likelihood (MLE) is one of the most often used function that fits this task of finding the set of parameters that best fit the observed data.

LLMs have three model architectures (a) encoder only (BERT) (b) decoder only (GPT) (c) encoder-decoder (T5). Looking at (b), which is a probability distribution over a word given the prompt, which is arrived by taking a smoothed exponential (softmax) of scores calculated using scaled dot products between each new prediction word with prompt, we use MLE to find the best distribution that fits the observed data.

Stochastic gradient descent (SGD) and ADAM (adaptive moment estimation) are two common methods used to optimize the objective function. The latter is memory intensive. There are many knobs like size of a floating point, calculating moments (more value per parameter), changing learning rates among others that can be used to learn a model. Sometimes the knob settings result in generic learning and other times overfitting. More often than not we just don’t converge i.e. no learning. ADAM is popular optimizer (I use it on the open source transformers from hugging), it keeps 3X more values per parameter than vanilla SGD. AdaFactor is an optimization on ADAM to reduce memory consumption but has been known to not always work.

A rule of thumb in ML is gather more data as a dumb model with lots of data beats a smart model with limited data. But training on large amount of data is costly. It used computational resources and as the whole process is iterative, we need fast processors to crunch through the data so the model can be adapted and iterated upon. More data does not guarantee convergence i.e learning. The whole exercise in learning a model looks and feels like art bordering on black magic than anything analytic or scientific. If modeling felt like a recipe than this is like cooking the recipe. The end result has a lot of variance.  

Saturday, June 24, 2023

The Model behind LLM and Transformers

 We want to get to a point where we have a probability distribution over a sequence of tokens. But what are these tokens? They are the words in a sentence that are quantized i.e. they are numbers. These numbers are not 64bit floats, they are more like 16 or 8 bit floats. In an array, these numbers map to a word and some of the context of the word. A classic example of a word (embedding) context is King - Man = Queen. This is a legit operation in this space. So we take a sentence and convert them into tokens and create vectors for each word which as a whole is called word embeddings. This seems like preliminary stuff, but it is not, this is so important and we get this wrong, the whole generative AI stuff falls on its face.

The engine of a generative AI car is the transformer. The main parts of the transformer is the encoder and decoder. Encoder is where you say “Yo quiero tacobell” and decoder translates to “I love Tacobell”. As an aside, if you tokenize as “I love Taco Bell” that will result a vastly different generation than if you tokenize as “I love Tacobell”. But to make this generative i.e. change its task model, they got rid of the encoder and used stacks of decoders. After all, given a prompt, we need to generate an essay where every word is picked from a probability distribution. A stack of decoders helps this more than a encoder/decoder architecture. This transformer architecture has two main components. First is the positional encoding and second is the multi-head attention.

If we start with positional encoding, we are looking at a matrix where each row is a word in the sentence of the prompt. This matrix is created by using 18th century math which showed us that a summation of sine and cosine function with varying frequencies carries all the information in the input signal. The input signal here is a prompt and to give different emphasis to the word position in the prompt, we use sine and cosines function to arrive at the rows in the matrix. So far we started with word embeddings and now we have a matrix with positional embeddings. But why? Well, the output of each time-step is fed to the decoder. E.g. I love Tacobell around the corner where the italics means they were generated in the previous step, will give emphasis to the words the corner by over the first few words. This is the part about attention. That brings us to the next big component called Multi-Head Attention (MHA).

To understand MHA, we need to understand the concept of query (Q), key (K) and Value (V). Query is the question asked by the transformer to find the correct next word in the sentence. In our case, that would be “around”. So query is asking is “around” the best fit given the input sequence i.e the Keys. The values are the input sequence’s embeddings that we calculated earlier. At the end of the day, we are using matrix multiplication, specifically dot product to guage the relative fit of a word given the input sentence. As we generate more words, we want the generated words to be in the input to find the next best match. A dot product is literally telling us how far is the next word (my query) from my current sentence (keys).

Note that till now we have not really used a neutral network. We would need that because so far all we have done is linear transforms (matrix multiplication). We need something nonlinear to activate the neurons or else what’s the point of calling this NN? That part is the feed forward neural network which gets this output (a matrix) that we calculated. But by doing all this processing and in parallel because we use multiple heads of attention, we are able to accelerate it. This was the original intent of transformers i.e. accelerate the translation using GPUs. If you are student of history, this is kind of like old mechanical engineering techniques to perform switching and routing. They worked but eventually were replaced by electronics. I think AI is lacking that paradigm shift. It hasn’t yet found its transistor.

Saturday, June 10, 2023

Perplexity, Entropy: How to measure LLMs?

 How can we measure efficacy of a language model? Language model researchers use the term “Perplexity” to measure how a language model performs on tasks on standard datasets. In a language model the task means quizzing the model to complete a sentence or hold a Q&A or generate an essay. GPT-3 scored well, in fact very well, on perplexity on standard benchmarks like Penn Tree Bank. Overall, though, the results were mediocre.

Perplexity of a model measures the “surprise” factor in the generation or how many branches does the model deal with when predicting the next word. A perplexity of 20 would mean that given a few words the model has to pick between 20 choices for the next word. If that number was 2, then the model has an easier task but that is most likely because we overfit the model to a specific task. Without this context specific training, the GPT-3 folks claim that the model is a few shot learner, which means it takes a few attempts before the model hones in on the task’s context. There are variants of transformer models like BERT cased - which are trained for specific tasks like “complete last word” or “fill in the blank” which perform much better at those specific tasks than GPT-3.

As we are building LLMs to perform task like translation, generation, completion etc., should we overfit a model to a specific task or leave it as a generic model and provide in-context training via prompting? With the GPT-3, it seems prompting is the chosen route to get the model to hone-in. And what about the model size? Does it make sense to overfit a model with hundreds of billions of parameters to do a single task (like translate) or leave it as a few shot performer i.e. mediocre? These are the tradeoffs that one has to make when building a new LLM. Larger models are mediocre at best.

Sunday, May 28, 2023

Language Models

 LM are probability distribution over sequences of tokens - which are words in a vocabulary and so the sequence is a phrase. Each phrase fragment gets a probability. The probability is higher for fragments which are good phrases. Good meaning grammatically correct and semantically plausible. Good is very dependent on extra information that is not in the training set. For example, a sentence “I saw the golden gate bridge flying into san francisco” needs some contextual information like bridges are stationary and “I” refers to a human in the plane. This means the LM needs to determine semantic plausibility of a phrase fragment. This is what makes LMs deceptively simple and easy to get wrong.

Mathematically, the phrase fragment is called sequence of tokens. And the tokens are pulled from a set “V” for vocabulary of the language. Probability distribution of a four token sequence assigns different probabilities to the ordering of the four tokens in that sequence. Some orderings are semantically implausible and others are syntactically incorrect. So far, we haven’t done any generation, but that is also possible in the LM, where given a sequence of tokens, say four, we can pull five token sequences and judge on its “goodness”. The judgement is based on some parameter which could be level starting from poor to best. To continue on the sentence above, the probabilities of “I saw .. sf on saturday”, “I saw ..sf yesterday” are all good as they are good sentences, but we need a new parameter to pick one over the other. This parameter is the randomness of search. If we are strict, we would pick the highest next probability of each word in the sentence and its next prediction. If we are lenient, we would randomize the next pick. This conditional generation is then determined by the prefix sequence, in our case “I saw .. sf” and the completion will be “yesterday”. This is the big problem with the LM, a lenient generation will create absurd sentences and even total fabrications and strict generation will create templates. Add to that there is a heavy reliance on the prefix sequence or “prompt”. This parameter is called temperature and ranges from lenient to strict and controls the variability of the generation. So, we now have the mathematical vocabulary of generation. We have a prompt, a completion and temperature.

A slight change to the prompt will generate a entirely different sentence. A longer prompt also changes the generation and any foreign words in the prompt, perfectly normal is spoken english, can also influence the generation in unpredictable ways. This length of the number of tokens in the prompt is a big determinate of computational complexity of LMs. A generation that uses all the tokens, such as RNN (Recurrent Neural Network) will take a very long time to train. A transformer which sets attention to a few prior tokens and exploits parallelism of GPUs limits the accuracy of the generation. Parallelism of transformers enable large LMs. Large as in trillion parameters. What was observed in later 2022 was the larger the model got, the fixed length generation coupled with temperature generated sentences which never existed before. This is called emergent behavior - this is beginning of “learning”. Repeated prompts on the same topic teach the LM to hone in on the context. This makes the completions sound more like spoken language and less like an automaton. This observation of emergent behavior without having to change the model (like no new gradient descents) is what is causing most the hype around Generative AI. As the model hones into a context, it feels like a turing machine where the generation feels conversational.

The biggest challenge for Gen AI - as an industry - is not to create larger and larger models, but instead to slice and dice an existing LLM into a size that can be deployed at scale. We can see the sizes of large model in this paper, reproduced below. It is rumored GPT-4 is over 1 trillion parameters as well. Training costs around $28K per 1 billion parameters. It is however a one-time cost. The continual cost is on the sliced/diced version - perhaps running on the phone - which needs to cost no more a penny per 1000 generations.

The models also need to be trained on proprietary and current data for them to generate economic value beyond research.

Expert Systems to Generative AI — tiny steps that caused giant leaps in productivity

In the beginning, we wrote the rules by hand. The expert systems of the 1980s and 90s were magnificent and exhausting. You found the best pe...