Today ConfigIQ sizes one model at a time. Sizing a set of models (e.g. for an architecture doc or a customer deployment plan) means running the flow once per model and manually collecting the results into a table.
Request: Let a user submit a list of models and get sizing back for all of them in one run.
Proposed shape:
- Input a list of models (paste HuggingFace IDs, or multi-select from the model picker), plus one shared workload profile (GPU, ISL, OSL, TTFT/concurrency target, precision).
- ConfigIQ runs the recommend flow for each model and returns a single results table: model, GPUs/replica (TP), replicas for target, memory/replica, TTFT, TPOT, throughput.
- Export the table to CSV.
- Handle unsupported models gracefully in the output (mark "not supported in AIC" rather than failing the whole batch).
Why: Real sizing tasks are almost always multi-model (a customer runs a fleet, an arch doc lists 10+ models). One-at-a-time is slow and error-prone when assembling comparisons.
Notes:
- Per-model workload overrides could come later; a shared workload profile covers the common case.
- Pairs well with the existing CSV export on the performance page.
Today ConfigIQ sizes one model at a time. Sizing a set of models (e.g. for an architecture doc or a customer deployment plan) means running the flow once per model and manually collecting the results into a table.
Request: Let a user submit a list of models and get sizing back for all of them in one run.
Proposed shape:
Why: Real sizing tasks are almost always multi-model (a customer runs a fleet, an arch doc lists 10+ models). One-at-a-time is slow and error-prone when assembling comparisons.
Notes: