Published event
ArtificialIntelligence ModelRelease 1 source(s)

Google Cloud C4 Brings a 70% TCO improvement on GPT OSS with Intel and Hugging Face

Updated September 26, 2026 · 2:45 PM · source date October 16, 2025

Summary

Google Cloud C4 Brings a 70% TCO improvement on GPT OSS with Intel and Hugging Face Google Cloud C4 Brings a 70% TCO improvement on GPT OSS with Intel and Hugging Face Published October 16, 2025 Update on GitHub Upvote 18 Jiqing.Feng Jiqing Intel Matrix Yao MatrixYao Intel Ke Ding kding1 Intel Ilyas Moutawwakil IlyasMoutawwakil Intel and Hugging Face collaborated to demonstrate the real-world value of upgrading to Google’s latest C4 Virtual Machine (VM) running on Intel® Xeon® 6 processors (codenamed Granite Rapids (GNR)). We specifically wanted to benchmark improvements in the text generation performance of OpenAI GPT OSS Large Language Model(LLM).

Why it matters

This ModelRelease is relevant to the technology intelligence record because it involves Google, Intel, Hugging Face, GitHub. The source article should remain the factual reference for follow-up coverage.

Key facts
  • Google Cloud C4 Brings a 70% TCO improvement on GPT OSS with Intel and Hugging Face Published October 16, 2025 Update on GitHub Upvote 18 Jiqing.Feng Jiqing Intel Matrix Yao MatrixYao Intel Ke Ding kding1 Intel Ilyas Moutawwakil IlyasMoutawwakil Intel and Hugging Face collaborated to demonstrate the real-world value of upgrading to Google’s latest C4 Virtual Machine (VM) running on Intel® Xeon® 6 processors (codenamed Granite Rapids (GNR)).
  • We specifically wanted to benchmark improvements in the text generation performance of OpenAI GPT OSS Large Language Model(LLM).
  • The results are in, and they are impressive, demonstrating a 1.7x improvement in Total Cost of Ownership(TCO) over the previous-generation Google C3 VM instances.
  • The Google Cloud C4 VM instance further resulted in: 1.4x to 1.7x TPOT throughput/vCPU/dollar Lower price per hour over C3 VM Introduction GPT OSS is a common name for an open-source Mixture of Experts (MoE) model released by OpenAI.
  • An MoE model is a deep neural network architecture that uses specialized “expert” sub-networks and a “gating network” to decide which experts to use for a given input.
  • MoE models allow you to scale your model capacity efficiently without linearly scaling compute costs.
Entities in this story
Related events