
Game Server Orchestration
Game server orchestration is the automated allocation, scaling and teardown of game server instances as matches begin and end. It decides where each server runs, starts it, hands players its address, and reclaims it when the match is over. Doing this by hand stops being practical once a game runs more than a handful of matches at once.
Also called
orchestration
,
What Does Game Server Orchestration Do for One Match?
Most studios arrive here with a simpler question: how do I host my game's servers? Orchestration answers the part that repeats, once per match, many times a day.
A server is requested. When players are ready to play, your backend calls the orchestrator's API for a server. What makes the call does not matter: a matchmaker, a lobby or your own service. If a suitable server is already running, players can join it instead.
A location is chosen. The orchestrator decides where the server runs, from what it knows about the players and the capacity available.
A server starts. It launches a container from the game's server image, usually a headless build of the dedicated server.
Players get the address. The orchestrator returns an IP and port, and your backend passes them to the players.
The match ends, the server stops. Its capacity goes back for the next match.
Repeated for every match, those steps add up to the job orchestration exists for: scaling with player demand. It adds servers as players arrive and removes them as they leave, both within one location and across many, so every match a studio's game needs gets a server, ideally at a location where the match plays well online. In infrastructure terms that is horizontal scaling, adding more servers of the same size, rather than vertical scaling, giving one server a bigger machine. Game servers rarely scale vertically: a match's needs are fixed by its build, so capacity grows by adding servers.
Two decisions shape how an orchestrator does this: where servers are allowed to run, and how that capacity is paid for.
Where Should Servers Run, and How Is Capacity Paid For?
Where: regional hubs or close to each match. Many games run servers in a few large server regions, and players are sent to the nearest one. It is cost-effective for the studio: capacity is concentrated, easy to operate, and each region holds a large pool of players to match. The cost lands on the players farthest from a hub, who play every match with the extra distance. The alternative is to choose a location for each match from many smaller ones, based on where that match's players are. It narrows the distance for more players, and it narrows the gap between players in the same match, which keeps the playing field level. That matters in a casual match as much as a competitive one: a match that feels fair is usually a match players enjoy more. It also asks more of the orchestrator, which has to make that choice in real time.
How: fixed or cloud capacity.
Fixed capacity is bare metal or reserved hosts, rented by the month. It costs less per hour when busy, often includes egress, and suits load that never drops: persistent worlds, or the baseline of matches running at every hour. Player demand is rarely flat, though. It typically peaks from late afternoon into the evening, local time, and falls overnight, so a host sized for the peak runs partly empty for the rest of the day. You pay for it while it sits idle, and it only exists in the locations you rented. That makes fixed capacity a return-on-investment decision: how many hours a day it is busy enough to cost less than paying per use.
Cloud capacity is started on demand and billed while it runs, either as your own cloud fleet or as multi-tenant capacity shared with other studios. Sharing spreads the cost of many locations across many games, so each one can reach more places than it would rent alone.
Why Do Most Games End Up Hybrid?
Because most games have both kinds of load. A match-based game runs a steady baseline of matches around the clock, and peaks on top of it: evenings, weekends, a launch, a content update. Fixed capacity is the cost-effective home for the baseline. Cloud capacity can absorb the peaks and reach players far from the fixed hosts. Running both under one orchestrator is hybrid orchestration, and the right mix differs for every game.
The edge cases sit at either end. Fortnite runs "almost completely" on AWS, its game server fleet included (DatacenterDynamics, February 2023). Persistent MMOs lean the other way, because a world that never shuts down is the definition of load that never drops: The Isle runs on Edgegap's Private Fleet, and Path of Titans has too.
On Edgegap, the two sides are Private Fleet, with 58 locations, and Edge Cloud, with 615+ locations across 17+ providers, under one orchestrator (platform data, 18 September 2026). To see where the line between fixed and cloud falls for your own concurrency, model it in the pricing calculator.
What Decides the Right Model for Your Game?
The game server, more than the orchestrator. For most games the answer is a balance of cost-effectiveness and player experience, and four properties of the server set where that balance lands:
Session length. Matches that end can move between locations and capacity types freely. A persistent world needs a home that stays up.
Resources per match. CPU and memory per match, measured in a profiled build, set how many matches fit on a host, and so what a fixed host is worth. See vCPU.
Players per server. A 1v1 is placed for two people; a 100-player lobby is placed for a crowd spread across a map.
Matches per process. A multi-room server packs many matches into one process, which changes the unit you size and pay for.
Start from those, then pick the building blocks: containers and server images for packaging, auto-scaling and warm pools for capacity, fleet management for fixed hosts. Whether to build this yourself is covered under game server hosting.
Those properties change as a game grows, launches in new regions or adds a persistent mode, so the orchestrator should let the mix change with them, without lock-in: standard containers, fixed and cloud capacity side by side, and cloud billed by the minute. Edgegap is built that way.
A word from our sponsor (ourselves!)
Most studios find their scaling ceiling at launch, in front of players. Edgegap reached 14 million concurrent sessions in a one-hour benchmark on real game server containers, so the limit you plan around is your game's, not your host's.
Place for the Player, Then Price It
Regional hubs are a cost-effective way to run servers. The players far from them pay the difference in latency, every match.
We measured what placement is worth. In a replay of a AAA publisher's live traffic, deploying each session's dedicated server at the location best suited to its players cut average round-trip time by 58%, from 116 ms to 49 ms, and put 78% of players under 50 ms, against 14% on the publisher's public-cloud setup (case study, 2019). Players notice. In a 2022 survey of 2,000 players in the US and UK, 1 in 3 said they stop playing when latency hits, and 42% said they would play more without it (connectivity report, July 2022).
So we start from the player: place every match where its players are, then make that cost-effective with the mix, a fixed baseline with cloud for everything above it. It is how Edgegap has orchestrated 135 million game server deployments since February 2019 (platform data, 18 September 2026).
Edgegap's Take (just our opinion, take it with a grain of salt!)
Does game server orchestration work with Unity, Unreal and Godot?
Yes. An orchestrator runs a container image, so it does not depend on the engine that built the server. The engine matters at two points: building a Linux server for the image, and a small integration so the server can read its match data and report when it is ready or finished. Unity, Unreal and Godot servers all fit this pattern, as do servers written without an engine, such as a Node.js server for a web game.
,









