Blog · Hosting

Crusoe Expands AI Platform Beyond GPU Rentals with Fine-Tuning and Inference Tools

Crusoe is broadening its managed AI platform with two new services, serverless fine-tuning and self-service inference deployments, aimed at enterprises that want to customize open-weight AI models without directly managing GPU infrastructure. The additions arrive through the company’s Intelligence Foundry platform and let customers fine-tune open-source foundation models, push them to managed inference endpoints, or export the resulting weights for use elsewhere.

The launch reflects a shift in how AI infrastructure providers compete. According to Dave McCarthy, research vice president at IDC, GPU availability alone is no longer the deciding factor for enterprise buyers.

  • “GPU access was the story for about 18 months. It’s not anymore, or at least it’s not the whole story,” McCarthy said, noting that fine-tuning pipelines, evaluation, deployment tooling, and inference optimization now matter more.
  • McCarthy warned that providers selling only raw compute risk becoming commoditized, while those managing the full model lifecycle, from training data through production monitoring, will differentiate.
  • He added that portability of models across platforms has become “a procurement requirement” rather than a nice-to-have.

Erwan Menard, senior vice president of product at Crusoe, said open-weight models have “crossed the quality threshold,” pushing enterprises to want ownership of their own model versions rather than depending on third-party APIs that can change underneath them. He said demand for continuous fine-tuning, where production data is regularly fed back into open-weight models, is accelerating faster than expected, particularly among teams building production AI agents that require predictable model behavior and clear data ownership.

The platform currently supports a curated set of open-weight models, including Qwen, DeepSeek, Gemma, and GPT-OSS.

Rather than requiring reserved GPU clusters, Crusoe schedules fine-tuning jobs dynamically across its infrastructure, since Menard said fine-tuning workloads tend to be spiky and reserved capacity forces teams to overbuy. The service restarts interrupted jobs automatically, saves training checkpoints, and stops billing once a model stops improving. Finished models are delivered in the open .safetensors format so customers can deploy them on Crusoe or elsewhere.

For production inference, the new Self-Serve Deployments service runs on Nvidia H100 and H200 GPUs, letting customers stand up managed inference endpoints without provisioning hardware themselves.

Serverless Fine-Tuning and Self-Serve Deployments are scheduled to become generally available next week through Crusoe Intelligence Foundry. Fine-tuning will be priced per million tokens, while inference deployments will be billed per GPU hour.