Free Quiz
Write for Us
Learn Artificial Intelligence and Machine Learning
  • Artificial Intelligence
  • Data Science
    • Language R
    • Deep Learning
    • Tableau
  • Machine Learning
  • Python
  • Blockchain
  • Crypto
  • Big Data
  • NFT
  • Technology
  • Interview Questions
  • Others
    • News
    • Startups
    • Books
  • Artificial Intelligence
  • Data Science
    • Language R
    • Deep Learning
    • Tableau
  • Machine Learning
  • Python
  • Blockchain
  • Crypto
  • Big Data
  • NFT
  • Technology
  • Interview Questions
  • Others
    • News
    • Startups
    • Books
Learn Artificial Intelligence and Machine Learning
No Result
View All Result

Home » DeepSeek launch ‘sparse attention’ model that cuts API costs in half

DeepSeek launch ‘sparse attention’ model that cuts API costs in half

Tarun Khanna by Tarun Khanna
September 30, 2025
in Artificial Intelligence
Reading Time: 2 mins read
0
DeepSeek launch ‘sparse attention’ model that cuts API costs in half

Photo Credit: https://techcrunch.com/

Share on FacebookShare on TwitterShare on LinkedInShare on WhatsApp

Researchers at DeepSeek on Monday launched a new experimental model known as V3.2-exp, designed to have dramatically decrease inference prices when used in long-context operations. DeepSeek introduced the model with a post on Hugging Face, also posting a linked academic paper on GitHub.

The most improtant feature of the brand new model is referred to as DeepSeek Sparse Attention, an complicated system defined in detail in the diagram below. In essence, the system makes use of a module referred to as a “lightning indexer” to prioritize unique excerpts from the context window. After that, a separate system referred to as a “fine-grained token choice system” chooses unique tokens from the ones excerpts to load into the module’s limited attention window. Taken collectively, they permit the Sparse Attention models to function over long quantities portions of context with relatively small server loads.

Photo Credit: https://techcrunch.com/

For long-context operations, the advantages of the system are significant. Preliminary testing by using DeepSeek determined that the price of a simple API call can be decreased by as much as a lot as half in long-context conditions. Further testing out may be needed to build a more robust assessment, however due to the fact the model is open-weight and freely to be had on Hugging Face, it won’t be long before third-party tests can assess the claims made within the paper.

Also Read:

Brazil releases AI supercomputer push, splits projects between Chinese, US companies

Nvidia just showed that the harness, not the AI model, is now the real hero

OpenAI to lease huge new AI data center in US, backed by Nvidia

Google packs Search and Gemini with new AI study tools

DeepSeek’s new model is one among a string of latest breakthroughs tackling the trouble of inference costs— importantly, the server costs of operating a pre-trained AI model, as distinct from the cost of training it. In DeepSeek’s case, the researchers have been seeking out ways to make the essential transformer architecture operate more efficiently — and finding that there are great improvements to be made.

Based in China, DeepSeek has been an unusual figure in the AI boom, especially for those who view AI research as a nationalist battle between the U.S. And China. The company made waves at the beginning of the year with its R1 model, trained the usage of mainly reinforcement learning at a far lower value than its American competition. But the model has no longer sparked a wholesale revolution in AI training, as some anticipated, and the company has receded from the highlight in the months since.

The new “sparse attention” method is not likely to provide the same uproar as R1 — but it can still teach U.S. Vendors some much wished tricks to assist keep inference costs low.

ShareTweetShareSend
Previous Post

Engineers generate Soft Robots That Can Literally Walk on Water

Next Post

Fiscal Fears Fuel Flight to Bitcoin, Gold as Major Currencies Falter

Tarun Khanna

Tarun Khanna

Founder DeepTech Bytes - Data Scientist | Author | IT Consultant
Tarun Khanna is a versatile and accomplished Data Scientist, with expertise in IT Consultancy as well as Specialization in Software Development and Digital Marketing Solutions.

Related Posts

OpenAI slows advanced AI development after cyberattack
Artificial Intelligence

OpenAI slows advanced AI development after cyberattack

August 19, 2026
Apple Builds China-Specific AI Model With Alibaba Support
Artificial Intelligence

Apple Builds China-Specific AI Model With Alibaba Support

August 19, 2026
Why Applied AI Engineering is Replacing Traditional Model Training in 2026
Artificial Intelligence

Why Applied AI Engineering is Replacing Traditional Model Training in 2026

August 14, 2026
NVIDIA Releases Nemotron 3.5 Lightning and NeMo Switchyard for More Efficient Agentic AI
Artificial Intelligence

NVIDIA Releases Nemotron 3.5 Lightning and NeMo Switchyard for More Efficient Agentic AI

August 14, 2026
Next Post
Fiscal Fears Fuel Flight to Bitcoin, Gold as Major Currencies Falter

Fiscal Fears Fuel Flight to Bitcoin, Gold as Major Currencies Falter

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

92 − = 90

TRENDING

How can Artificial Intelligence Maximize Your Business Growth in 2021?

artificial intelligence
by Tarun Khanna
March 18, 2021
0
ShareTweetShareSend

This AI mines the numbers buried in scientific papers and turns them into usable data fast

This AI mines the numbers buried in scientific papers and turns them into usable data fast

Image Credit: https://techxplore.com/

by Tarun Khanna
April 22, 2026
0
ShareTweetShareSend

Toward a latest framework to accelerate large language model inference

Toward a latest framework to accelerate large language model inference

Schematic diagram of SPECTRA and other existing training-free approaches. Photo Credit: https://techxplore.com/

by Tarun Khanna
August 8, 2025
0
ShareTweetShareSend

An unreleased Anthropic model made progress on one of math’s biggest unsolved problems

An unreleased Anthropic model made progress on one of math’s biggest unsolved problems

Image Credit: https://techcrunch.com/

by Tarun Khanna
August 12, 2026
0
ShareTweetShareSend

Researchers broaden ‘SyMerge’ technology maximizing AI model synergy

Researchers broaden ‘SyMerge’ technology maximizing AI model synergy

Principles and effects of the AI model merging technology (SyMerge) developed by Professor Sung-Eun Hong's research team at Sungkyunkwan University. (Left) Process of AI knowledge merging and compatibility measurement, (Right) Performance effect achieved by adapting just a single core layer. Image Credit: https://techxplore.com/

by Tarun Khanna
July 17, 2026
0
ShareTweetShareSend

SpaceX’s near-term AI payoff seen anchored to Earth, not outer space

SpaceX’s near-term AI payoff seen anchored to Earth, not outer space

Image Credit: https://www.reuters.com/

by Tarun Khanna
July 13, 2026
0
ShareTweetShareSend

DeepTech Bytes

Deep Tech Bytes is a global standard digital zine that brings multiple facets of deep technology including Artificial Intelligence (AI), Machine Learning (ML), Data Science, Blockchain, Robotics,Python, Big Data, Deep Learning and more.
Deep Tech Bytes on Google News

Quick Links

  • Home
  • Affiliate Programs
  • About Us
  • Write For Us
  • Submit Startup Story
  • Advertise With Us
  • Terms of Service
  • Disclaimer
  • Cookies Policy
  • Privacy Policy
  • DMCA
  • Contact Us

Topics

  • Artificial Intelligence
  • Data Science
  • Python
  • Machine Learning
  • Deep Learning
  • Big Data
  • Blockchain
  • Tableau
  • Cryptocurrency
  • NFT
  • Technology
  • News
  • Startups
  • Books
  • Interview Questions

Connect

For PR Agencies & Content Writers:

connect@deeptechbytes.com

Facebook Twitter Linkedin Instagram
Listen on Apple Podcasts
Listen on Google Podcasts
Listen on Google Podcasts
Listen on Google Podcasts
DMCA.com Protection Status

© 2024 Designed by AK Network Solutions

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Artificial Intelligence
  • Data Science
    • Language R
    • Deep Learning
    • Tableau
  • Machine Learning
  • Python
  • Blockchain
  • Crypto
  • Big Data
  • NFT
  • Technology
  • Interview Questions
  • Others
    • News
    • Startups
    • Books

© 2023. Designed by AK Network Solutions