Pruna AI Releases Open-Source Framework for AI Model Optimization
Pruna AI, a European startup specializing in AI model compression, has made its optimization framework open source. The framework integrates various efficiency methods like caching, pruning, quantization, and distillation, helping developers compress AI models while maintaining performance. The company aims to standardize these optimization techniques in a manner similar to how Hugging Face standardized transformers and diffusers.
Big AI labs have used similar techniques in-house, such as OpenAI’s distillation approach for GPT-4 Turbo. However, Pruna AI’s framework uniquely combines multiple methods, making them accessible and easy to implement. Currently, the company is focusing on optimizing image and video generation models, with clients such as Scenario and PhotoRoom.
In addition to the free open-source version, Pruna AI offers a paid enterprise solution featuring an automated optimization agent. This agent allows developers to specify performance goals, and it automatically finds the best compression configuration. Pruna AI monetizes its pro version with a pricing model similar to GPU rentals on cloud services, promising significant cost savings in AI inference.
The startup recently secured $6.5 million in seed funding from investors including EQT Ventures, Daphni, Motier Ventures, and Kima Ventures, further supporting its development efforts.
