โŒ

Reading view

Cloudflare Tests Cache Transcoding to Reduce Storage Requirements

Cloudflare recently described a prototype called Cache Transcoding that compresses eligible cache content, mainly uncompressed text such as HTML, JSON, CSS, and JavaScript, using Zstandard before storing it on disk. The hyperscaler estimates that the approach could provide petabytes of additional effective cache capacity, although broader testing is still needed.

By Renato Losio
  •  

Gemma 4 Multi-Token Prediction Delivers up to ~3x Faster Token Generation

Gemma 4 can be paired with multi-token prediction (MTP) drafters that use speculative decoding to generate multiple tokens in parallel, allowing the model to verify them in a single pass and achieve up to ~3รƒโ€” faster inference without quality loss.

By Sergio De Simone
  •  
โŒ