Different models fail the same discipline differently.
A strong model may over-comply and conceal whether the protocol mattered. A weak model may drown in the architecture. Another may preserve the posture but lose a critical gate under context pressure. These are different safety findings.