Deploy Llm Presets
Dedicated Models
Deploy Llm Presets
DeepInfra presets and mirrored vLLM recipes for hf_repo_id, told apart by
source; empty when none. Filter by gpu/engine/source.
GET
Deploy Llm Presets
Authorizations
Bearer authentication header of the form Bearer <token>, where <token> is your auth token.
Query Parameters
Available options:
L4-24GB, L40S-48GB, A100-80GB, H100-80GB, H200-141GB, B200-180GB, B300-270GB, RTXPRO6000-96GB, other Response
Successful Response
Preset id.
Allowed Nx configs.
Config source.
Inference engine.
Engine tuning knobs.
Raw engine flags; vLLM recipes only.
Short display name (e.g. "Throughput-optimized").