AMD Revolutionizes AI Inference with Taalas Acquisition, Unlocking 16,000 Tokens Per Second
AMD's acquisition of Taalas, a Canadian AI startup, brings a groundbreaking technology that embeds AI models directly into silicon, resulting in unprecedented inference speeds of over 16,000 tokens per second. This move is set to disrupt the AI landscape, offering developers and businesses a significant boost in performance and efficiency.
AMD is buying Canadian startup Taalas, which hard-codes model weights directly into inference chips. That makes them extremely fast but locks each chip to a single model. A demo chip hit over 16,000 tokens per second per user running Llama 3.1-8B. Google is reportedly working on a similar approach for Gemini. The article AMD acquires Taalas, a startup that bakes AI models directly into silicon appeared first on The Decoder.