At the late peak, certification was slow, login and page-turning, but it was not even possible to get a page on it. The most classic complaints scene in Portal. Users often reacted badly to the system, but the real reason was that little was a single failure or more than one layer. We dealt with such problems as first step not by adding machines, but rather by stratification, because Cadens were different from other layers, but they changed their methods.
Level 1: Authentication of server performance
The most top layer is the authentication server itself. Evening peaks and requests for certification are pressed up, and single servers can only line up if they do not have a cluster, CPU and connections reach the top. We've seen schools carry the whole school with an entry server, it's okay during the day, 9 o'clock to log in. This level of judgement is simple, and there are signs one hour before the peak, depending on the server CPU and active connection curve.
Second tier: database bottlenecks
The authentication checks the account numbers and writes logs, which end up in a database. If you focus on one library, late peak disks and locks will rise sharply, as evidenced by extended and occasional failures in certification. We have processed a school with an application server but a database disk runs full and authenticated with a sample card. This layer is to be defused by literacy separation, hot spot data caches and log strides, and light-viewing application layers indicators are not found.
Level 3: Webback
Portal authentication messages often go around the central machine room certification server, and there are delays in the network return process across buildings if there is a congestion or route. We measured that the pre-certification device was sunk to the area, and the delay could be reduced from dozens of milliseconds to several milliseconds, with an immediate difference in login sense. This layer of problems can be seen by end-to-end connectivity testing and application layer grab packages, not just by looking at servers.
Level 4: Access devices
The most easy to ignore is the access switch and wireless access point itself. The old CPUs are full, they don't transfer authentication reports out of the house, and it doesn't help if they are top-heavy. We have seen a building that changes its authentication server or card, and we find that the floor switches are too old to transmit.
Layer 5: Portal Page Loading Dependence
Login itself loads resources, and if it relies on external addresses to solve slow, the user sees that page unopening. This pit is hidden because authentication logic does not collapse, but simply adds a placard to a bug, which can be misconstrued as certification failure. We usually ask Portal to localize its resources, to rely less on the outer web, and whether logs can open in seconds, directly affecting users ' perception of the whole system.
Positioning method: stratological pressure
When we do capacity assessments, we take individual pressure on each layer: pressure-certified interfaces look at the application layer, database-checking the storage layer, network-pressure, and access to the device layer. Any ring that delays surges with a sudden increase, the problem is that without this stratification, all amplification is blind guess.
Common error: server only
The most typical error is that one card plus the server, and the money is still in question because the bottlenecks are on the database or access level. We have seen the reverse, when access equipment should have been replaced but all the time, the servers were stacked together.
Cache and Fail
The architecture-level solution is the Cache for Hotspot Certification with a sharp limit on currents. Cache of high frequency certified status, reducing double checkups; priority for core links at peaks, and delay in non-critical actions. These two moves can smooth out the late peaks.
Shut up: Carton's not a magnifying solution.
The surveillance has to be stratified.
In order to be fast-tracked, we have to bury every layer of indicators. We give schools a watchboard showing the application level response, database loads, delay in return, access equipment health, and portal loading time. Which rings are unusual at first glance. Without stratification, peaks are difficult to guess, and the location is slow, and users start to curse early.
Precautionary actions in front of the peaks.
Carton is mostly preventable. We suggest that we check server connections, database disks, and access equipment CPUs one hour before the late peak, and find that there are not enough pre-flow limits or magnifications. It is much easier to save than a peak before it comes. Students feel better, and they can pick up a lot of calls at night.
Summary: Portal peak Carton is rarely a single cause, but it has to be located first on the specific level. Servers, databases, return journeys, accesses, portal loadings, five layers of tubes and repairs. Blind extension is just throwing money into water, so stratification is the right way.