A single race condition that can lock up the entire system? Wow, that's rough...
ORBITER-FORUM will be temporarily closed at 2026-07-23 18:00 UTC while we complete some OF maintenance tasks. The amount of downtime is expected to take up to one hour, but probably less.
A single race condition that can lock up the entire system? Wow, that's rough...
As long as only 2 jobs per second arrived, all was fine. But at 100 jobs per second, it took just 5 hours until the system is blocked.
And from what you're describing, I assume that's not due to lack of power, but simply because some jobs are never dequeued and pile up. Yeah, there should really always be a ttl on jobs waiting for a lock.
*thinks of code he wrote during the last month*
Uhm... If you'll excuse me, I just realised there's something I needed to do... :leaving:
You will have to explain all above in slow language before I even start smiling!
N.
I'm imagining something similar is going on in your database races...?
N.
Our problem was self inflicted, never mind the human element.
I'm imagining something similar is going on in your database races...?
N.
The end is near.
Somewhere in the most uninhabited reaches of the Andromeda Galaxy, underneath the outback of some uninhabited desert planet, or at any rate, far away from Wolfsburg, a German middle manager is digging with inhuman speed towards the core of the planet, crawling into the deepest, darkest hole he can until the storm passes.
And the good thing is, nobody will ever really know about the project outside its organisational context. No press, no angry users, no torches and forks. Even if maybe half of the planet is affected when the system failed, even the experienced users did not even know that this software was existing. It was hidden deep inside the backend of many other services, with only very few people knowing about its function. Just imagine that it took a small bug to find out about all connected systems: We accidentially corrected a wrong return code and suddenly found all systems, that directly connected to the webservice and had fixed the bug on their end.
Well, yeah, I went for the absolute worst-case scenario to drive home the point about what happens when unserviced jobs start to accumulate in memory until the server thrashes itself to death on swap. I'm not sure every step in that chain has ever happened all together, but certainly large chunks of it have. But still, even if the public doesn't know what you do, if the failure of your product leads to somebody having a PR disaster, even if that organization doesn't know who you are or what you do, they'll give their vendors grief, who will in turn give their vendors grief, until somebody who knows what you do gives you grief (and you get to bring out the 2 year old ticket). Crap, after hitting the proverbial fan, always rolls down the proverbial hill.
So... How do yousomething up as good as Game of thrones season 8? Like, even if you were trying, how?