The Quiet Land Grab in ML Infrastructure
Baseten has been doing something Modal’s growth team should find uncomfortable: converting ML engineers who came for serverless GPU access and staying for something far more specific – a production inference layer that actually handles the messy reality of deploying large models at scale.

What Baseten Is Building That Modal Isn’t
Modal built its reputation on developer experience. Clean Python decorators, fast cold starts, a serverless abstraction that made spinning up GPU workloads feel almost trivial. For ML experimentation, batch jobs, and prototype pipelines, it remains genuinely well-designed. The problem is that experimentation is not production, and the gap between those two states is where Baseten has been quietly setting up camp.
Baseten’s core product is Truss, an open-source model packaging framework, combined with a managed inference platform built around what the company calls “model serving primitives.” That sounds dry, but the practical difference is significant. Where Modal asks developers to write infrastructure logic into their application code, Baseten treats the model serving configuration as a first-class artifact – something you version, test, and deploy independently of the rest of your stack. For teams moving a model from research to production API, that separation matters more than any cold start benchmark.
The inference infrastructure conversation has also shifted around hardware. A growing number of ML teams are not just running standard GPU workloads – they are managing multi-GPU inference for models like Llama 3, Mixtral, and various fine-tuned variants that require careful memory management across A100 or H100 clusters. Baseten has invested heavily in making that specific scenario work without requiring teams to become distributed systems engineers. Modal is capable of running these workloads, but it is a more manual configuration process, and manual configuration in production is where outages happen.
There is also the question of who Baseten is actually targeting. The company has been deliberate about going after mid-market ML teams inside companies that are past the proof-of-concept stage but not large enough to run a dedicated MLOps function. These teams want something opinionated. They do not want infinite flexibility – they want the right decision made for them on model routing, replica scaling, and request batching. Baseten’s product is increasingly built around that assumption, which makes it a different bet than Modal’s more generalist positioning.

Why Modal’s Dev Base Is Vulnerable
Modal attracted a particular kind of user: technically sharp, ML-focused, and slightly allergic to YAML-heavy infrastructure tooling. That is exactly the persona Baseten is now courting with its inference-specific feature set. When a tool purpose-built for your exact use case enters the market, the switching cost question becomes very real, especially when the incumbent was never specifically designed for that use case in the first place.
The inference layer is a stickier product than the compute layer. Once a team has configured model routing logic, batching strategies, and autoscaling rules inside Baseten’s platform, that configuration represents institutional knowledge. It gets documented, reviewed, and eventually owned by someone whose job title includes “ML platform.” That is a very different retention dynamic than a developer running batch jobs who can switch tools between projects. Baseten is building toward that stickiness deliberately, and it shows in the product roadmap emphasis on observability, A/B testing for model versions, and latency-tier routing.
Modal’s response to the inference trend has been incremental rather than strategic. The platform has added GPU support, improved cold start times, and expanded its function-level caching options – all genuine improvements, but none of them address the core gap, which is that Modal still requires developers to think like infrastructure engineers when they want production-grade model serving. Baseten’s pitch is that you should not have to. That message is landing with the teams who tried Modal for inference and found themselves writing more boilerplate than they expected.
Pricing is a secondary factor, but worth noting without inventing numbers. Both platforms charge for GPU compute, but Baseten’s production tier includes managed features – autoscaling, request queuing, health monitoring – that Modal users typically have to build themselves or pay for separately through third-party tooling. For a team that was already planning to buy those capabilities somewhere, Baseten often represents a simpler total cost calculation even if the headline compute rate is comparable.
The developer community signal is subtle but directional. On forums where ML engineers share infrastructure frustrations, Baseten mentions have increased alongside discussions about moving models to production. That kind of organic presence is hard to manufacture and usually precedes a larger adoption wave. Modal still dominates conversations about serverless GPU experimentation, but Baseten is increasingly the name that comes up when the conversation turns to “we need this to actually work in production.”

The Inference Market Is Not a Single Bet
It would be a mistake to frame this as a zero-sum fight. ML teams often run both platforms at different stages of their workflow – Modal for fast iteration, Baseten for what gets shipped. That split creates a natural entry point for Baseten to consolidate the relationship over time, because the production workload is where the budget actually lives. Experimentation tools are often expensed on a credit card; production inference is a line item in a quarterly budget review.
What makes Baseten’s position interesting is not that it is outrunning Modal on general developer adoption – it is not. Modal still has more name recognition in the broader developer tooling space, and its community is active. The tension is specifically in the inference segment, where Baseten has a product advantage that is difficult to neutralize without Modal making a product pivot that cuts against its own positioning. Whether Modal treats that segment as worth defending or accepts a natural specialization in the market is the question that will define the next twelve months for both companies.









