StiCAS: Compositional Activation Steering via Stiefel Manifold Coordinate Transport
StiCAS learns orthogonal concept subspaces via Stiefel optimization so compositional activation steering achieves exact commutativity and doubles multi-concept success.
Published 2026Sydney Poster Session 3 · Wed, Dec 9, 10:00 AM–1:00 PM local time · Hall 1-4OpenReview ↗
Only vote on papers you've read. Sign in with GitHub to vote.
Abstract
Activation steering is a lightweight way to control pretrained generative models: a learned map inserted into the model's activations can steer generation toward target behavior without updating the model weights. While effective for single target behavior, existing steering methods remain brittle when multiple behaviors are targeted. A style intervention followed by a safety intervention can produce a different result than the same applied in the reverse order, as the second intervention is evaluated on activations already shifted by the first. We trace this failure to coordinate sharing: existing transport-based steering methods learn different concepts in the same activation coordinates, so their interventions interfere by construction. To address this, we introduce StiCAS (Stifefel Compositional Activation Steering), a geometric steering framework that learns a separate orthogonal subspace for each concept. Each concept is steered by an affine transport restricted to its own subspace, and orthogonality between subspaces makes these transports non-interfering by construction. The concept frame is trained directly on the Stiefel manifold, so Riemannian optimization preserves the orthogonality required for composition throughout learning. This structure gives an exact commutativity guarantee: any sequential ordering of single-concept interventions reproduces the simultaneous multi-concept update. Across 45 concept pairs and three open models, StiCAS drives order-dependent interference exactly to zero and roughly doubles joint-concept success over the strongest activation-transport baseline.