How should each model run?
Set workload requirements, latency limits, quality floors, and a trial budget. Let the specialists propose changes and use measured trials to decide what stays.
You get A history of tested configurations and their outcomes.