Why does storing monitoring metrics in a relational database like MySQL become a bottleneck?
The write volume is huge, and reading a metric over a time range means pulling it out of a table row by row, which gets slow as history grows.
Writing: 4000 systems with 200 values each, polled every 15 seconds, is 800,000 values per poll round, more than 3 million rows a minute, all to be inserted into a table.
Reading is worse. Trending never asks for one value; it asks "how did CPU use develop over the past year?" The software then fires queries at the table and reads the matching rows one by one, transposing them into a series. Anyone who has asked PNP4Nagios (a Nagios plug-in that graphs performance data from a relational backend) for a year of web-server load knows how long that takes.
So sooner or later the database is guaranteed to become the performance choke point for trending. That problem is what time series databases were built to solve.
Go deeper:
InfluxData — Time series database explained — why time-stamped data overwhelms general-purpose databases.