Damn it! Specs of 4GB RAM Potatoes Can Run 70B Giant AI, The Secret to This Airllm Makes Shock!

How to run a local AI on a low spec PC using AirLLM. Save solution for computer users with limited memory.

The world of artificial intelligence (AI) is currently being crazy about giant language models (LLM) such as Llama 3 or Mixtral who have extraordinary intelligence. However, there is one great wall that hinders ordinary users: hardware needs. To run AI Super Smart sized 70 billion parameters (70B), you normally need a Sultan’s computer with VRAM or RAM above 140GB!

But what if you only have an old computer with 4GB RAM? Does the dream of tasting a smart local AI have to run aground? The answer: no more. Thanks to an open-source Python library named Airllm, computers with “medium” specifications can now do things that were once considered impossible.

Getting to know Airllm: Computer Savior Low Specifications

Many users think that the only way to run AI on an old computer is to use a small model that has been trimmed to the fullest (quoted). The problem is, small models are often less clever and often “hallusinate” when given a complicated task.

This is where Airllm comes with a new game map. AirLLM is an innovation that allows us Running Local AI on Low Spec PC without having to cut or damage the quality of the original model’s intelligence. This project breaks the old law that “big models must use large GPUs”.

How does it work? The Secret Behind Layer Sharding

Airllm does not do magic, but uses a very clever memory engineering technique called Layer Sharding or sequential separation of layers.

Imagine a giant AI model as a book Encyclopedia as thick as 80 chapters. Ordinary applications will try to insert the entire book into the RAM of your computer at once. If your RAM is only 4GB, the computer will obviously break down or experience Crash (out of memory).

Airllm works in a different way. He only reads one chapter (one layer transformer) From storage media (SSD/NVME), insert it into 4GB RAM/VRAM memory, process your data row, then immediately delete it from memory to be replaced with the next chapter. This process repeats in relay from the first layer to the last layer.

By breaking this physical limitations, the memory used on your computer is kept below the 4GB number, it can even be as small as 1.6GB for a 70b model!

AirLM’s real capabilities for everyday users

It must be admitted honestly, this relay system has one big drawback: speed. Because the computer must constantly read and delete data from storage to memory, the process of producing text becomes very slow. Airllm takes several tens of seconds to minutes just to utter one word.

Therefore, if you hope you can CHAT Instantly interactive like using chatgpt or gemini, AirLLM is not a suitable tool. However, for daily tasks that are “run and then just sleep”, the ability of this tool is extraordinary:

  • Summarizing super long documents: Do you have a draft of hundreds of pages or complicated legal documents? Just put it in the AirLLM script, leave your computer, and when you return, you’ll get a very accurate summary of high-end AI intelligence.
  • Contextual Language Translation: In contrast to ordinary translators, the large model AI understands culture and language style. You can use it to translate long writing drafts offline with very natural results.
  • Program code checking (debugging): For novice students or programmers, you can use this giant model locally to detect hidden errors on complex lines of code without the need for an internet connection.
  • Total data privacy: Because all processes run 100% locally on your hard drive, your sensitive data or confidential documents are guaranteed safe from third-party server snub.

Reality that must be known before trying

Although this technology sounds great, daily users must understand the performance limits. The most crucial component if you want to try AirLLM is not an expensive graphics card (GPU), but the speed of your storage media. Using a high speed NVMe SSD is required so that the transfer process between layers of the model does not take too long.

If your goal is to find a responsive, agile, and interactive AI assistant to accompany work Real-time On a 4GB RAM computer, switch to an application like Ollama or LM Studio Loading small quantized models (such as llama 3 8B or PHI-3) will be much more efficient and provide convenience for daily use.

However, if you are a researcher, a writer who works in an area without a signal, or technology enthusiastic who wants to test the maximum limit of your old computer, AirLLM is a clear proof that the limitations of hardware specifications are no longer the end of everything in the era of artificial intelligence.

Leave a Reply

Your email address will not be published. Required fields are marked *


Baca Juga

Back to top button

Adblock Detected

LidahTekno.com is supported by Google Adsense advertising to provide content for you.Please consider disabling AdBlocker or adding us to your whitelist so we can continue providing the best technology information and tips.Thank you for your support!