GauS models operator scheduling via Gaussian reparameterization to capture time's ordinal nature, cutting optimization space and yielding Pareto-optimal results.
Flash-KMeans eliminates GPU HBM bottlenecks via fused assignment and inverse mapping updates, delivering up to 17.9x speedups over existing exact k-means implementations.
TNQE uses structured unitary tensor networks to learn shallow, resource-efficient quantum data encoding circuits that achieve 0.04x the depth of amplitude encoding and scale to high-resolution images on real hardware.
SIMBA is a GPU-accelerated MBA synthesizer using cache-free bottom-up enumeration to scale beyond prior CPU and cache-based GPU tools. It achieves substantial speedups, handles larger specifications, and solves expression sizes existing methods cannot.