❌

Normal view

  • ✇MIT Technology Review
  • Understanding the thermal ceiling in portable power Shuo Yang
    Plug a phone into a modern charger and the first 10 minutes are impressive. The next 20 are not. This is not a defect. It’s the connected device protecting itself. As temperature rises during charging, a smartphone’s battery management system reduces the current it will accept, because heat accelerates the chemical degradation that permanently reduces battery capacity. The charger may be capable of delivering more, but the device simply stops taking it. For anyone building products in the
     

Understanding the thermal ceiling in portable power

9 September 2026 at 16:18

Plug a phone into a modern charger and the first 10 minutes are impressive. The next 20 are not.

This is not a defect. It’s the connected device protecting itself. As temperature rises during charging, a smartphone’s battery management system reduces the current it will accept, because heat accelerates the chemical degradation that permanently reduces battery capacity. The charger may be capable of delivering more, but the device simply stops taking it.

For anyone building products in the portable power category, this creates an uncomfortable gap between specification and experience. A device rated at 25 watts is accurate in the sense that it can deliver 25 watts. Whether it delivers 25 watts for the duration of a charge is a different question, and one the specification does not answer.

The specification gap

The gap matters commercially because it is invisible at the point of purchase and obvious in use.

Consumers compare wattage figures on packaging. They don’t compare thermal curves, because thermal curves are not published publicly. The result is a category where products differentiate on a number that describes peak output rather than sustained output, and where the actual user experience of two products with identical specifications can diverge substantially.

This is particularly acute in magnetic wireless charging. Inductive power transfer generates heat at both the transmitting and receiving coils, and the magnetic attachment that makes these products convenient also places the heat source in direct contact with the device it is charging. Convenience and thermal performance are working against each other by design.

The industry’s response for the past several years has been materials science. Graphite sheets, thermal interface materials, conductive housings, and heat-spreading layers have all improved how efficiently accumulated heat moves away from the source. Each generation has been incrementally better than the last.

But passive dissipation has a structural limitation: it can only move heat that has already been generated, and only as fast as the surrounding air will accept it. In a sealed, pocket-sized enclosure, that ceiling arrives quickly. Improving the materials slows the rate of temperature rise. It does not prevent the temperature rise.

Moving from dissipation to removal

The alternative is active thermal management, which is standard in stationary electronics and largely absent from portable ones for reasons that are easy to understand. Fans add volume, weight, moving parts, and noise. In a product category defined by portability, each of those is a meaningful cost.

At Anker, which manufactures charging and power products, engineering teams spent the past several development cycles working on whether that tradeoff could be made acceptable rather than eliminated. The approach involves several interacting systems: a micro centrifugal fan, dual airflow channels routed to avoid interference with the magnetic array, a three-layer graphene heat-spreading layer, and a control algorithm that modulates fan speed based on real-time temperature and battery state rather than running at a fixed rate. The result is that the Anker MagGo Power Bank 2 Pro has become the world’s fastest and coolest wireless power bank.

In internal testing, at 77 °F (25 °C) ambient, the back of the power bank stays below 96.8 °F (36 °C) throughout wireless charging, 21.6 °F (12 °C) below the international standard limit of 118.4 °F (48 °C), for a comfortable grip. Comparable magnetic power banks in the same testing typically reached 113 °F (45 °C) or higher within 20 minutes. The functional consequence is that the connected device does not reach the threshold at which it begins reducing charge acceptance, so 25 watts of Qi2.2 magnetic wireless charging is delivered as a working rate rather than an opening rate. In practice, an iPhone 17 Pro reaches 50% charge in 25 minutes. The Anker MagGo Power Bank 2 Pro’s premium performance in both charging speed and thermal management is certified by SGS, an independent testing and certification company.

The same principle applies in reverse. Recharging a power bank generates heat too, which is why devices in this category are often slow to recharge, leaving users with an empty accessory at the moment they need it. Active cooling during input allows the unit to accept 45 watts and reach 80% in 52 minutes.

What this suggests about the category

There is a broader pattern here worth naming, because it is not unique to charging.

When a category improves along a single axis for long enough, the constraint usually migrates somewhere else. Charging spent a decade optimizing power delivery. Power delivery is now, for most practical purposes, solved: the electronics can supply more energy than the receiving device is willing to accept. The binding constraint moved to thermal management, and the industry continued optimizing the axis it had always optimized, because that is the axis the specifications describe.

Recognizing when a constraint has moved is difficult precisely because the old metric keeps improving. Wattage figures have continued to climb. Products have continued to get faster on paper. The measurement stayed valid while quietly ceasing to describe the thing users experience.

For product organizations, the practical question is whether their specifications still measure the constraint or merely measure the capability. The two align until the constraint shifts and specifications rarely shift with it.

The transparency problem

A second implication follows from the first. If sustained performance differs meaningfully from peak performance, and if only peak performance is disclosed, then buyers cannot evaluate the products in front of them.

This is one reason Anker is adding displays on charging products. The Anker MagGo Power Bank 2 Pro shows real-time power, temperature, battery level, and estimated time remaining. Some of that is user convenience. But some of it is a Anker stating a deliberate position—this category deserves to have the complete and accurate data made transparent to all.

Anker expects independent reviewers to test these claims and considers our internal numbers to be the correct outcome. The gap between specification and experience closes faster when the experience is measurable. The Anker MagGo Power Bank 2 Pro will be available in the U.S. on September 17, 2026.

This content was produced by Anker. It was not written by MIT Technology Review’s editorial staff.



vAttention: Verified Sparse Attention

arXiv:2510.05688v2 Announce Type: replace-cross Abstract: State-of-the-art sparse attention methods for reducing decoding latency fall into two main categories: approximate top-$k$ (and its extension, top-$p$) and recently introduced sampling-based estimation. However, these approaches are fundamentally limited in their ability to approximate full attention: they fail to provide consistent approximations across heads and query vectors and, most critically, lack guarantees on approximation quality, limiting their practical deployment. We observe that top-$k$ and random sampling are complementary: top-$k$ performs well when attention scores are dominated by a few tokens, whereas random sampling provides better estimates when attention scores are relatively uniform. Building on this insight and leveraging the statistical guarantees of sampling, we introduce vAttention, the first practical sparse attention mechanism with user-specified $(\epsilon, \delta)$ guarantees on approximation accuracy (thus, "verified"). These guarantees make vAttention a compelling step toward practical, reliable deployment of sparse attention at scale. By unifying top-$k$ and sampling, vAttention outperforms both individually, delivering a superior quality-efficiency trade-off. Our experiments show that vAttention significantly improves the quality of sparse attention (e.g., $\sim$4.5 percentage points for Llama 3.1 8B Instruct and DeepSeek-R1-Distill-Llama-8B on RULER-HARD), and effectively bridges the gap between full and sparse attention (e.g., across datasets, it matches full model quality with up to 20x sparsity). We also demonstrate that it can be deployed in reasoning scenarios to achieve fast decoding without compromising model quality (e.g., vAttention achieves full model quality on AIME2024 at 10x sparsity with up to 32K token generations). Code: https://github.com/skylight-org/sparse-attention-hub. Webpage: https://sky-light.eecs.berkeley.edu.

Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference

arXiv:2511.16449v5 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have shown great potential for embodied AI by integrating visual perception, language understanding, and action execution. In real-time deployment, these models must process continuous visual streams, incurring substantial computational overhead. Visual token pruning -- a mainstream technique for accelerating Vision-Language Models (VLMs) by retaining salient tokens while discarding redundant ones -- offers a natural candidate solution to this challenge. However, directly applying VLM-oriented pruning methods to VLA inference can cause severe degradation in manipulation performance. Our analysis attributes this degradation to a key mismatch: VLA inference exhibits distinct attention patterns between the vision-language prefill stage and the action-decode stage, so pruning based only on context-prefill semantic salience is biased toward semantic cues and may remove action-critical visual tokens. Motivated by this observation, we propose VLA-Pruner, an effective plug-and-play token pruning method grounded in the visual requirements of VLA inference, further exploiting the temporal continuity of robot manipulation. Specifically, VLA-Pruner estimates visual-token importance from both semantic prefilling and temporally smoothed action relevance, and then applies a Combine-then-Filter strategy to retain compact, non-redundant tokens under the compute budget. Experiments show that VLA-Pruner outperforms state-of-the-art approaches across multiple VLA architectures, achieving up to 1.99x speedup with comparable manipulation quality.

Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs

arXiv:2603.22446v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has significantly improved reasoning in large language models (LLMs), yet the token-level mechanisms underlying these improvements remain unclear. We present a systematic empirical study of RLVR's distributional effects organized around three main analyses: (1) token-level characterization of distributional shifts between base and RL models, (2) the impact of token-level distributional shifts on sequence-level reasoning performance through cross-sampling interventions, and (3) fine-grained mechanics of these shifts at the token level. We find that RL fine-tuning induces highly sparse and targeted changes, with only a small fraction of token distributions exhibiting meaningful divergence between the base and RL policies. We further characterize the structure and evolution of these shifts through analyses of token entropy, positional concentration, and reallocation of probability mass. To assess the functional importance of these sparse changes, we conduct cross-sampling experiments that selectively swap token choices between the base and RL models with varying intervention budgets. We show that inserting only a small fraction of RL-sampled tokens into base generations progressively recovers RL performance gains, while injecting a similarly small number of base token choices into otherwise RL-generated sequences collapses performance to base levels, isolating a small set of token-level decisions directly responsible for RLVR's performance gains. Finally, we explore divergence-weighted variants of the advantage signal as a diagnostic intervention, finding that they can yield improvements over baselines. Together, our results shed light on the distributional changes induced by RLVR and provide a fine-grained, token-level lens for understanding RLVR fine-tuning as a targeted refinement process.

ToolTree: Efficient LLM Agent Tool Planning via Dual-Feedback Monte Carlo Tree Search and Bidirectional Pruning

arXiv:2603.12740v1 Announce Type: new Abstract: Large Language Model (LLM) agents are increasingly applied to complex, multi-step tasks that require interaction with diverse external tools across various domains. However, current LLM agent tool planning methods typically rely on greedy, reactive tool selection strategies that lack foresight and fail to account for inter-tool dependencies. In this paper, we present ToolTree, a novel Monte Carlo tree search-inspired planning paradigm for tool planning. ToolTree explores possible tool usage trajectories using a dual-stage LLM evaluation and bidirectional pruning mechanism that enables the agent to make informed, adaptive decisions over extended tool-use sequences while pruning less promising branches before and after the tool execution. Empirical evaluations across both open-set and closed-set tool planning tasks on 4 benchmarks demonstrate that ToolTree consistently improves performance while keeping the highest efficiency, achieving an average gain of around 10\% compared to the state-of-the-art planning paradigm.
❌