Why does nobody talk about rate limits being the last line of defense when everything else fails
I used to think rate limiting was just a boring throttle you set and forget. Then last March I was on a call with a team in Austin after their API got hammered for 6 hours straight. Their auth was fine, their tokens were scoped right, but a partner integration leaked a key into a public repo and someone wrote a script that pulled 2.1 million customer records in one night. The only thing that slowed it down was a rate limit of 100 requests per minute per key, which bought their on call guy enough time to rotate everything. I was on that call and I watched the graph go from a flat line to a wall. Now I push back hard when people say rate limits are just for cost control. They are not. They are the seatbelt you never notice until the crash. If your API has no per key and per IP limits with alerts tied to them, you are one leaked key away from a very bad morning. What limits do you actually run in prod, and did you ever test them with a real load script?
Wait, 2.1 million records in one night with just auth and scoping in place? That is terrifying and honestly makes me feel dumb for never testing my own limits with a real load script. I always assumed a leaked key would get caught by scoping, but clearly that is not enough if nobody is watching the request volume. Going to go check my per key limits this week.