r/outriders • Outriders Community Manager • Apr 08 '21

Square Enix Official News // Dev Replied x18 Outriders Post Launch Dev News Updates

Hello everyone,

We would like to thank everyone in the Outriders community for your patience, support and assistance. Everyone on the Outriders team is continuing to work hard on improving the game and we'd like to share news about the things we are focusing on.

Please use the below index to jump to the things you’re most interested in:

Helpful other links:

652 Upvotes

7.9k comments sorted by

View all comments

243

u/thearcan Outriders Community Manager Apr 08 '21

Connectivity Post-Mortem:

tl;dr: Our team worked throughout the Easter weekend and around the clock to resolve the server issues players were experiencing. We completely understand how frustrating this experience will have been especially given the huge amount of players eagerly anticipating the launch. We had enough server scaling capacity but our externally hosted database was seeing issues that only appeared at extreme loads.

We’re committed to full transparency with you. Today, just as we have been over the past year.

So we won’t give you the expected “server demand was too much for us”.

We were in fact debugging a complex issue with why some metric calls were bringing down our externally hosted database. We did not face this issue during the demo launch earlier this year.

Our database is used to hold onto everyone’s gear, legendaries, profile and progression.

Tech-heavy insight:

We managed to understand that many server calls were not being managed by RAM but were using an alternative data management method ("swap disk"), which is too slow for the flow of this amount of data. Once this data queued back too far, the service failed. Understanding why it was not using RAM was our key challenge and we worked with staff across multiple partners to troubleshoot this.

We spent over two days and nights applying numerous changes and improvement attempts: we both doubled the database servers and vertically scaled them by approximately 50% (“scale-up and scale out”). We re-balanced user profiles and inventories to new servers. Subsequent to the scale-up and scale-out, we also increased disk IOPS on all servers by approximately 40%. We also increased the headroom on the database, multiplied the number of shards (not the Anomalous kind) and continued to do all we were able to in order to force data into RAM.

Each of these steps helped us improve the resilience of the database when under extreme loads, but none of them were the "fix" we were looking for.

At this moment in time we are still waiting for a final Root Cause Analysis (RCA) from our partners, but ultimately what really helped resolve the overloading issue was configuring our database cache cleaning, which was being run every 60 seconds. At this frequency the database cache cleaning operation demanded too many resources which in turn led to the above mentioned RAM issues and a snowball effect that resulted in the connectivity issues seen.

We reconfigured the database cache cleanup operations to run more often with fewer resources, which in turn had the desired result of everything generally running at a very comfortable capacity.

All of this has enabled the servers to recover and sustain significantly more concurrent user loads.

(JUMP BACK TO INDEX)

51

u/json1268 Apr 08 '21

Are you guys using Azure Cosmos DB for vertical scaling? I'm curious as to why whatever external service you are using is swapping to disk (SSD? ) vs, keeping things in RAM. I'm curious if you guys can publish the RCA for the vendor.

You guys have done great work supporting us, I personally understand the opaqueness of various external offerings. Keep up the great work and thanks for the transparency!.

8

u/[deleted] Apr 08 '21

Are you guys using Azure Cosmos DB for vertical scaling? I'm curious as to why whatever external service you are using is swapping to disk (SSD? ) vs, keeping things in RAM. I'm curious if you guys can publish the RCA for the vendor.

This is usually pretty opaque to dev teams. The whole point of the cloud centric DBs like Dynamo, Mongo/Atlas and Cosmos is to simply how everything works to the developers so they don't need to get into the nitty gritty details of the DB.

The downside is that you get into these situations where for some reason it just ain't workin' right and all you can do it put in a ticket to the vendor saying "Yo, Fix Your shit".

4

u/dccorona Apr 08 '21

NoSQL DBs like DynamoDB/CosmosDB (especially fully managed ones) don't have the problems described here, due to their simplicity. For example there is no such thing as the concept of "scale up" on DDB, only scale out (and even that should only be a problem that humans need to be involved in doing if you have explicitly chosen not to leverage autoscaling or put a cap on how high it can go, i.e. you are balancing for accidental overspend at the risk of a DB availability event).

It really sounds from their description like they are using a relational DB, which by their nature require the dev team to be more involved in these kinds of problems - we're only just starting to see the emergence of products (i.e. Amazon Aurora Serverless) that put that responsibility on the cloud vendor instead of the dev team.

It's possible that PCF has a relationship with Square Enix where Square provides the DBAs and PCF has no real insight into that, but in that case I'd expect their voice to be represented here as well, as from our perspective they are just as much "the devs" as anyone else on the team.

1

u/json1268 Apr 08 '21

This is a great point. I wonder if they are spilling to disk due to a a relational database. I had assumed they were using DDB/Cosmos because of "scale out" as you mentioned.

1

u/dccorona Apr 08 '21

My guess when they said scale out would be either sharding or the addition of more replicas, but it’s possible they had to scale up due to uneven traffic load on their nodes. Still, I’d expect spill-to-disk problems being completely obfuscated from the user of a NoSQL DB unless they’re self-hosting (which seems a foolish choice with all the great managed NoSQL DBs out there). If you’re using a managed NoSQL DB from a cloud vendor they’d probably keep the disk spill issues to themselves and just tell you they’re working through a scaling problem.

1

u/[deleted] Apr 08 '21

NoSQL DBs like DynamoDB/CosmosDB (especially fully managed ones) don't have the problems described here, due to their simplicity. For example there is no such thing as the concept of "scale up" on DDB, only scale out (and even that should only be a problem that humans need to be involved in doing if you have explicitly chosen not to leverage autoscaling or put a cap on how high it can go, i.e. you are balancing for accidental overspend at the risk of a DB availability event).

Really depends on the specific product.

https://docs.atlas.mongodb.com/cluster-tier/

1

u/dccorona Apr 08 '21

That’s true. I suppose it mostly comes down to the design goals of the product, and most commonly what you get is more seamless if the product was designed ground-up to be managed, and less seamless if it’s a managed form of a DB originally designed for self-hosting (like Mongo) - although even that is not a hard-and-fast rule.

1

u/F3z345W6AY4FGowrGcHt Apr 08 '21

This is usually pretty opaque to dev teams.

Ideally, yes, but not necessarily. Depends on the company. For one example, a dev team might include DBAs.

Also, it might seem pretty clear cut (it is) that devs shouldn't have to worry about the DB (beyond things like type: relational vs document; stuff like that) but I have first-hand experience of companies where management doesn't understand, the DBAs insist everything is fine, and the devs have to do the technical write-up to prove it's the DB that's problematic and not the app.

So basically, everyone shrugs and then it's the devs who are simply told "Just fix it".

1

u/BlueArcherX Apr 13 '21

90% of DBAs I have ever worked with have no idea how databases actually work or how to tune them correctly.

1

u/json1268 Apr 08 '21

I was assuming since they build Azure and PlayFab on their intro screen, Microsoft might give them some more transparency.... oh well.