Overclocking: Diagnose Before You Push
A practical method for PC instability: test memory, read SMART data, run a burn-in, and know what each result actually proves before you overclock.

Why does a stress test pass while the machine still crashes?
An unstable PC is a diagnostic problem before it is a tuning problem. Run the machine at stock settings first, isolate the failing subsystem with targeted tests, and only then change one variable at a time. If you skip that order, you are guessing, and guessing costs more hours than testing. Because a stress test only proves that the workload it ran did not fail during the time it ran. That is a narrow claim, and it is easy to misread as a general statement about stability. Most stress tools load one or two subsystems hard: CPU math, memory bandwidth, GPU shaders, or a mix. A crash that happens when the machine is idle, or when a specific driver loads, or when the disk wakes from sleep, sits outside that envelope. The test passes, the machine still falls over, and the user concludes the test is useless. It is not useless. It answered a different question. The practical fix is to match the test to the symptom. Random reboots under light load point at memory, power delivery, or a marginal idle voltage. Crashes only under sustained all-core load point at thermals, VRM behavior, or an unstable core offset. Crashes during file copies or game asset streaming point at storage, the memory controller, or the PCIe link. Write down what the machine was doing in the ten seconds before it failed. That log is worth more than any single benchmark run.
What does a burn-in actually prove?
A burn-in proves that the system survived a defined load for a defined duration under a defined ambient temperature. Nothing more, and that is already useful. Three variables define the test: the workload, the duration, and the environment. Change any one and you have run a different test. A thirty-minute run in a cold room tells you less than a two-hour run in a warm one, because heat soak is often what pushes a marginal system over the edge. The first ten minutes mostly measure the cooler. The next hour measures the whole loop: paste, mounting pressure, case airflow, VRM heatsink, and the fan curve that ties them together. Duration matters for a second reason. Some instability is intermittent. A memory error that appears once per few billion transfers will not show up in a short run, and it will absolutely show up during a long render or a compile. If your workload is long, your test should be long. If your workload is short and bursty, a long burn-in is less relevant than a rapid cycle test that loads and unloads the system repeatedly. Read the failure, not just the pass or fail. A hard power-off is a different signal from a blue screen, which is different from a worker thread reporting a wrong result, which is different from a thermal throttle that quietly drops clocks. Each one points somewhere else. A wrong computation with no crash is the classic memory or cache signature. A sudden power-off under load is the classic power delivery signature. A crash that only happens after twenty minutes is the classic heat soak signature.
How do you read an instability without chasing ghosts?
Change one variable per run, and keep a written record. That single habit eliminates most wasted evenings. Start from a known state. Load defaults, disable any overclock, and confirm the machine is stable at stock. If it is not stable at stock, you have a hardware or firmware problem, not a tuning problem, and no amount of voltage tuning will fix it cleanly. This is where a lot of people lose days: they tune around a fault that was there before they started. Then introduce one change. A memory frequency step, a core offset, a fan curve, a BIOS update. Run the same test you ran before, for the same duration, in the same room. Compare. If the result changed, you learned something. If it did not, revert and try the next single change. Two habits make this faster. First, note ambient temperature with every run, because a result from a cold morning is not comparable to one from a warm afternoon. Second, note the exact failure mode, not just that it failed. "Crashed" is not data. "Worker 7 reported a rounding error after 41 minutes at 78 C" is data. Expect diminishing returns. At some point the extra frequency costs more voltage, more heat, and more risk than the performance is worth. The honest question is not how far the chip can go. It is whether the gain shows up in the work you actually do. For most workloads, a stable machine at a modest setting beats a fast machine that fails once a week.
What should you check before touching any setting?
Check the boring things first, because they cause a large share of reported instability. Memory seating and slot population. A module not fully clicked into its slot can pass a short test and fail under thermal expansion. Consult the board manual for the correct slots when running two modules, since the wrong pair can force a lower stable frequency. Cooler mounting and paste. Uneven mounting pressure produces a hot core that shows up only under sustained load. Compare per-core temperatures, not just the package number. A single core running ten degrees hotter than its neighbors is a mounting or paste signal. Power supply headroom and cabling. Transient spikes from a modern CPU or GPU can exceed what a marginal unit handles, and the failure looks like a random shutdown rather than a clean error. Use separate cables rather than daisy-chained connectors where the card allows it. Firmware and drivers. A BIOS release note that mentions memory compatibility is worth reading. So is a chipset driver update. These are cheap changes with real effects, and they are much easier to revert than a voltage curve. Storage health. A drive with reallocated sectors or a rising error count can produce crashes that look like memory faults. Read the attributes, understand which ones are predictive and which are not, and replace a drive that is clearly degrading before you spend a weekend tuning around it.
When is the machine actually stable enough?
Stable enough means it survives your real workload, repeatedly, with margin left over. Define your real workload precisely. If you compile large projects, run a compile. If you render, render. If you game for three hours, do that. Synthetic tests are efficient proxies, but the final check is the work itself, because the work has a load pattern no benchmark reproduces exactly. Leave margin. A setting that passes at the edge of thermal and voltage limits will fail in summer, or with dust in the heatsink, or with an extra drive installed. Back off one step from the edge. The performance difference is usually small and the reliability difference is not. Re-test after any change. A new GPU, a BIOS update, a moved case, a new fan: each one changes the thermal and power picture. The old result no longer applies. Finally, keep the record. A short log of settings, ambient temperature, test duration, and outcome turns a future problem into a twenty-minute check instead of a fresh investigation. That log is the actual deliverable of all this testing, more than any single stable profile.
A short method to reuse
Start at stock and confirm stability. Match the test to the symptom. Change one variable per run. Record ambient temperature and the exact failure mode. Back off one step from the edge. Re-test after every hardware or firmware change. If you want a reference for the diagnostic side, including how to build a bootable toolkit and read sensor data with its limits in mind, the material at overclockix.com covers that ground in more depth than a single article can.
For a structured way to separate these cases, including how to read SMART attributes and interpret voltage and temperature sensors with their individual limits, the method laid out at hardware stability testing is a reasonable starting framework. The point is not the tool list. The point is that each result has a boundary, and you should know where yours sits.
The same rule applies on an RF bench, and measuring before adjusting covers isolation, insertion loss and how to verify a switch.


