DiagnosticIQ benchmarks LLM recommendation of industrial maintenance actions from symbolic rules across 6,690 questions, finding frontier models match human experts but break under structural perturbation due to calibration failures rather than capability gaps.
A generic patch Transformer achieves state-of-the-art zero-shot time series forecasting via simple training, with scaling and data ablations isolating key performance drivers.