Friday 14 March 2025
The quest for affordable and decentralized artificial intelligence (AI) has long been a holy grail for tech enthusiasts. A recent study published by researchers at University of California, Berkeley, takes a significant step towards achieving this goal.
The researchers have developed a system called DeServe, which enables the efficient deployment of large language models (LLMs) on consumer-grade GPUs in a decentralized manner. This means that individuals with access to these devices can contribute their computing power to the collective effort of processing massive amounts of data, without having to rely on centralized cloud services.
The challenge lies in the fact that LLMs require enormous computational resources and high-speed networks to function efficiently. Traditionally, this has made it difficult for individuals to participate in AI-related activities, as they lack access to such infrastructure.
DeServe addresses this issue by introducing a novel serving system that leverages consumer-grade GPUs to accelerate LLM inference. The system consists of three key components: microbatches, KV cache offloading, and microbatch scheduling.
Microbatches refer to the process of dividing large tasks into smaller, manageable chunks. This allows DeServe to efficiently utilize available computing resources on individual devices, reducing the need for high-speed networks and centralized infrastructure.
KV cache offloading is a technique that enables DeServe to transfer unused cache memory from one device to another, freeing up space for more critical computations. This not only improves overall system performance but also reduces energy consumption.
Microbatch scheduling is responsible for optimizing the order in which microbatches are processed, ensuring that each device contributes its fair share of computational resources without overloading or underutilizing them.
The researchers tested DeServe using a variety of LLMs and found that it achieved significant improvements in throughput and efficiency compared to existing serving systems. The system was able to handle high-latency networks by adapting to network conditions and optimizing serving throughput accordingly.
DeServe has far-reaching implications for the future of AI development, as it opens up new possibilities for decentralized computing and collaboration. With this technology, individuals can contribute their computing power to larger projects, such as training and deploying LLMs, without relying on centralized infrastructure or cloud services.
The potential applications of DeServe are vast and varied. For instance, researchers could use the system to accelerate complex computations for scientific simulations, medical imaging, or climate modeling. Additionally, the technology could be applied in industries such as finance, healthcare, or education, where decentralized computing can enhance data processing and analysis capabilities.
Cite this article: “Decentralized AI: A Step Towards Affordable and Efficient Computing”, The Science Archive, 2025.
Artificial Intelligence, Decentralized Ai, Language Models, Gpus, Cloud Services, Consumer-Grade Devices, Microbatches, Kv Cache Offloading, Microbatch Scheduling, Serving Systems







