Complex AI models emit 50 times more CO₂
A study has shown that advanced AI language models can produce up to 50 times more CO₂ emissions compared to simpler alternatives, especially when generating lengthy responses. The energy efficiency and accuracy of these models vary significantly depending on their architecture and the type of tasks they are designed to perform.
Natura
Researchers from the University of Applied Sciences Hochschule München (Germany) conducted an analysis of the environmental impact of using large language models (LLMs). In their study, they tested 14 different LLMs—from basic to more advanced—using 1,000 identical standard questions. The number of tokens generated by each model was converted into greenhouse gas emissions.
Main Findings of the Study
The study found that models with advanced reasoning capabilities produce up to 50 times more CO₂ emissions compared to models that provide brief answers. This is because complex models generate more tokens and require significantly more computational power, which leads to higher energy consumption and, consequently, increased emissions.
To assess energy consumption, the researchers used a computer with an NVIDIA A100 GPU and the Perun framework, as well as an average emission factor of 480 g CO₂/kWh. Each of the 14 models answered 1,000 questions on topics such as philosophy, world history, international law, abstract algebra, and mathematics. The study included both text-based and reasoning models from companies like Meta, Alibaba, Deep Cognito, and Deepseek.
Model Comparison
On average, reasoning models generated 543.5 tokens per question, while text-based models produced about 37.7 tokens. However, generating more tokens did not always result in higher accuracy—often, the answers were simply more detailed.
- The most accurate model was Deep Cogito 70B (70 billion parameters) with an accuracy of 84.9%. However, it produced three times more emissions than similarly sized models that gave simpler answers.
- The most energy-intensive model was Deepseek R170B, generating 2,042 g CO₂-equivalent per 1,000 questions—comparable to a 15 km car trip. Its accuracy was 78.9%. According to calculations, if this model answered 600,000 questions, its emissions would be equivalent to a round-trip flight from London to New York.
- The most energy-efficient model was Alibaba Qwen 7B (27.7 g CO₂-equivalent), but its accuracy was only 31.9%.
Impact of Question Type
Energy consumption also depended on the type of question: tasks in abstract algebra and philosophy required more computation than simple questions. The study did not include some major LLMs, such as ChatGPT by OpenAI, Gemini by Google, Grok by X, and Claude by Anthropic.
The results of the study were published in the journal Frontiers in Communication.
