Analysis-driven transformer linearization isolates state update design to show delta-style networks outperform gated accumulation via key-dependent rank-1 projections, reducing approximation errors with sink tokens and cache routing to match adaptive caching at 32B scale.