How to Enable the Active Directory Recycle Bi
Accidentally deleting users, groups, or organizational ...




Two servers with identical specifications can perform very differently, and the reason is usually invisible in the specification sheet. The difference is not what your machine has, but how much of it you actually receive at the moment you need it — and on shared infrastructure, that quantity is not fixed. It depends on what the other tenants on the same physical host happen to be doing at that instant.
This is why benchmarks lie in a specific and predictable way. A fresh instance on a quiet host scores well. The same instance, unchanged, on the same plan, can score noticeably worse six weeks later when the host is busier. Nothing about your configuration changed. Your share of the hardware did.
Understanding this changes which questions are worth asking before you buy. Not "how fast is it", which has no stable answer, but "how much of the stated capacity will still be available to me during my busiest hour, on a host I do not control". That is a question about the allocation model, and it is answerable.
The assumption worth removing first is that a vCPU corresponds to a predictable slice of a physical core. It does not.
A vCPU is a virtual processor presented to your machine by the hypervisor, which is the software layer that schedules virtual processors onto real hardware. The mapping between vCPUs, physical cores, and hardware threads varies by provider, by processor architecture, and by the specific scheduling design in use. Two providers each selling "2 vCPU" may be selling different things, and the specification will not tell you which.
The mechanism that creates the variability is oversubscription: total virtual CPU capacity assigned across the machines on a host can exceed what that host can physically deliver if every one of them demanded full utilisation simultaneously. That is a normal and defensible engineering decision, because most workloads do not peak at the same time. It becomes a problem for you only when they do.
| Term | What it means | Why it matters to a buyer |
|---|---|---|
| vCPU | A virtual processor scheduled onto physical hardware by the hypervisor | Not a fixed fraction of a core — the ratio is provider-specific |
| Oversubscription | Assigned virtual capacity exceeds simultaneous physical capacity | Efficient by default; harmful when you are the one waiting |
| Contention | Multiple workloads needing the same physical CPU at once | The source of the variability you experience as slowness |
| Steal time | Time your vCPU was ready but the hypervisor ran someone else | The one guest-visible measurement of the above |
Steal time is the share of time your virtual CPU was ready to execute but could not, because the hypervisor was servicing another virtual processor. The definition is worth reading carefully, because it is narrower than it appears: it measures involuntary waiting at the virtualisation layer.
That makes it fundamentally different from the other figures people confuse it with.
The distinction is operational, not academic. High utilisation means your machine is working hard and you should look at your own workload. Steal time means your machine wanted to work and was not allowed to — and no amount of application tuning fixes that, because the constraint is outside your control entirely.
One practical warning about reading it. The counter in /proc/stat is cumulative — a total since boot, not a current rate. Reading it once gives you an average over the entire life of the machine, which will dilute a problem that only appears during your peak hours into something that looks negligible. Sampling over an interval is mandatory, not preferable, and it is the single most common reason people conclude they have no contention problem when they do.
The measurement itself is simple, and it is worth doing before you escalate anything to a provider, because a claim without samples is treated as an opinion.
| Command | What to read | Notes |
|---|---|---|
top |
The st value on the CPU line |
Fastest glance; per-host aggregate only |
vmstat 1 |
The st column |
One-second cadence, easy to watch over time |
mpstat -P ALL 5 60 |
The %steal column, per vCPU and aggregate |
The one to use for evidence — about five minutes of samples |
Run the third command during the period when the problem actually occurs, not when it is convenient. If your users complain at a specific hour, that is the window to sample, and the interval matters as much as the command.
Read the individual vCPU lines rather than only the aggregate. Averages conceal precisely the pattern that matters most: when a workload pins its busy workers to one or two CPUs, or when the scheduler concentrates runnable work onto a subset of them, a single vCPU can be starved repeatedly while the machine-wide average looks comfortable.
Some rough reference points are useful, with the caveat that they are heuristics rather than thresholds. Sustained steal under roughly 2% is generally healthy. Two to five percent suggests light contention worth watching. Five to ten percent indicates meaningful overselling risk. Sustained values above about ten percent indicate severe contention and justify evidence-gathering and escalation or migration. The caveat is real: there is no universal number that defines a problem, because a latency-sensitive API can be damaged by a level a nightly batch job would never notice. The right baseline is your own server's normal range, measured over time, compared against its behaviour now.
This is where most diagnoses go wrong, because three unrelated causes produce a similar symptom and only one of them is visible from inside your machine.
| Mechanism | What is actually happening | Where you can see it |
|---|---|---|
| Host contention (steal) | The hypervisor gave your vCPU time to another tenant | Inside the guest, via mpstat |
| Burst ceiling or credit exhaustion | You consumed your plan's allowance for elevated performance | Only in the provider's control panel |
| Internal cgroup or service limit | A quota inside your own machine capped one workload | Your own container or service configuration |
The middle case deserves emphasis because it is invisible from the guest. Some plans permit short bursts above a baseline and then hold you at the baseline once the allowance is consumed — a legitimate design that suits intermittent demand and surprises anyone who assumed the burst was the specification. Your steal time will be unremarkable throughout, because nothing is being withheld from you; you are being held to a level you already agreed to, defined in terms you may not have read. That behaviour has to be confirmed in the provider's console or documentation, not diagnosed from the server.
The third case produces a distinctive pattern: one service is slow while the rest of the machine is healthy. If that is what you observe, the constraint is a container or unit limit you configured, and the fix is in your own configuration rather than your provider's.
There is also a limit worth knowing about the first case. Steal time is a genuine signal where a platform exposes it, but it is not exhaustive — some virtualisation platforms have been documented to report zero steal under oversubscription in certain conditions. Use it when you have it, and treat it as one input rather than a verdict.
If contention is real and you intend to raise it with a provider, or to justify a migration, the difference between a complaint and a case is the shape of the evidence.
Correlation is the whole argument. Steal time on its own shows contention exists; it does not show that your workload is affected, and it certainly does not identify which neighbour or scheduling decision caused it. What makes it actionable is pairing the CPU samples with an application-level symptom measured over the same window — p95 latency, queue depth, request duration, or job completion time — so the two can be read together.
A decision table helps cut this short, because most apparent CPU problems are not contention at all.
| What you observe | Most likely cause | Right next step |
|---|---|---|
| High user and system time, low steal | Your own workload is consuming its allocation | Profile the application before touching infrastructure |
| Repeatable steal during latency spikes | Host-level contention is credible | Collect per-vCPU samples and escalate with correlation |
| High I/O wait, low steal | Storage latency, not CPU scheduling | Investigate disk behaviour and query patterns |
| Slower only after long bursts | A burst allowance was consumed | Check the provider's plan terms and console metrics |
| One service slow, host healthy | An internal quota is capping that workload | Review your own container or service limits |
The framing that makes this decision tractable is to stop asking which allocation model is faster and start asking which one matches the workload's shape.
Shared CPU is not a compromise when it fits. Development and staging environments, low-traffic sites, small internal APIs, scheduled background jobs, and queue consumers mostly wait on I/O or sit idle, and a well-matched shared plan runs them effectively for years. The efficiency of oversubscription is precisely why the entry price is low, and a workload that rarely needs sustained CPU collects that benefit without paying for it.
Shifting to dedicated allocation is justified when the requirement changes in kind rather than degree. Sustained computation such as encoding, image processing, or analysis. Latency-sensitive APIs where response-time consistency reaches users or downstream systems. Databases serving concurrent queries. Build runners, where unpredictable duration directly slows a team. Anything where a timing inconsistency is a product defect rather than a delay.
The distinction that matters is not average performance, which may be similar under light load, but predictability under sustained load. A shared instance frequently matches or exceeds a dedicated one on a quiet host. It diverges when the host is busy, and it diverges precisely when you most need it not to — which is why benchmarks run on fresh instances tell you so little.
There is a cost argument on both sides, and it does not always favour the cheaper option. Upgrading repeatedly to a larger shared instance to absorb variability you cannot control is a form of paying for the problem: bigger CPU and memory allowances to compensate for inconsistent delivery of the CPU you already bought. A smaller instance with predictable allocation can cost less in total than a larger unpredictable one, once the size increases are added up. The reverse also holds, and is more common — a mostly idle development server gains little from reserved capacity.
A single benchmark result on a new instance is nearly information-free. Two changes make it useful.
The first is to report workload-shaped results rather than abstract scores. A raw compute score is comparable across providers and tells you nothing about your application. What you want is the metric your users experience — p95 response time, requests per second at a fixed concurrency, sustained write throughput, job completion time — measured under conditions resembling production.
The second is to measure during the period you care about, and more than once. A test at 3am on a quiet host and a test at your busiest hour are different experiments, and the second is the one that predicts your experience. Run the same test repeatedly across different times and days; the variance between runs is itself a result, and often the most informative one you will get.
A practical rule for interpreting what you find: if latency rises together with your own CPU utilisation while steal stays near its baseline, the demand is yours and the answer is tuning or resizing. If latency rises together with sustained steal while your own utilisation is moderate, contention is the credible explanation and the numbers are worth escalating. If latency rises with I/O wait or disk queue depth, look at storage first. And if steal is elevated but latency and throughput are both fine, you have a capacity risk rather than a demonstrated problem — which is a reason to monitor rather than to migrate in a hurry.
Everything above converges on a different sizing method from the usual one.
Most purchase decisions are made against average utilisation, which is the number that least predicts whether a plan will hold up. The relevant figure is peak-hour behaviour under contention, and the relevant question is what happens to this workload when the host is busiest — not what it looks like on a quiet afternoon.
Four properties are worth verifying before committing, and none of them are performance benchmarks. Whether the stated CPU allocation is shared or reserved, and if shared, whether a burst allowance applies and what happens when it is exhausted. Whether the documentation states the oversubscription policy in terms you can act on. Whether steal time is exposed to the guest at all, since a platform that hides it removes your only direct measurement. And whether the plan can be changed without rebuilding the server, because the correct answer to this question often changes after you have run a real workload against it.
ULightHost is SurferCloud's entry tier, and the published comparison on its page is a useful illustration of why abstract scores mislead. Its Coremark results are given in both a single-core and a four-core form, sitting at 17,563 and 67,518 respectively, against 14,781 / 57,069, 14,329 / 53,485, and 9,227 / 33,649 for three well-known same-tier alternatives the page benchmarks against. Those figures say the entry tier is not slow, and they say nothing at all about whether the allocation is shared, how it behaves under contention, or what happens during a peak. That gap is not a criticism of the benchmark — it is what a compute score is for, and it is why the questions above matter more than the ranking. The tier is also built around pre-configured images and per-plan pricing at $2.9 per month for the 2 GB / 40 GB NVMe configuration, with longer-term terms available, across more than sixteen locations — practical for the low-demand workloads listed earlier, where predictable allocation matters less than cost.
Where the workload is sustained rather than intermittent, the comparison changes and the elastic compute range is the relevant tier, since it is the one built around reserved capacity rather than shared. If the deciding factor turns out to be latency consistency rather than throughput, the hourly options let you test the same workload across different hosts and hours before committing to a term — which is, given everything above, the most informative thing you can do with a small budget. The full ULightHost specifications and locations are published in detail.
The general principle transfers to any provider: buy the allocation model, not the clock speed. Two plans with identical specifications are not the same product if one reserves CPU and the other shares it, and the specification sheet will not tell you which is which.
What is CPU steal time?
The share of time your virtual CPU was ready to run but the hypervisor scheduled another virtual processor instead. It measures involuntary waiting at the virtualisation layer, not your own CPU usage.
What steal percentage is normal?
No universal figure exists. Under roughly 2% is generally healthy, 5 to 10% suggests meaningful overselling risk, and sustained values above about 10% indicate severe contention. What matters more is the deviation from your own server's established baseline.
Why is steal always near zero when I check?
Almost certainly because you are reading the cumulative counter in /proc/stat without sampling over an interval. Use mpstat with a defined interval and sample during the period when the problem actually occurs.
My server slows down after long bursts. Is that steal?
Probably not. That pattern points to a burst or credit allowance being exhausted, which is only visible in the provider's control panel and produces no steal time at all.
One of my services is slow but the machine looks fine — why?
Most likely a container or service quota inside your own machine. Check your cgroup or unit limits before investigating infrastructure.
Does dedicated CPU mean better performance?
Not necessarily faster — more predictable under sustained load. On a quiet host a shared instance can match or beat it. The difference appears when the host is busy, which is exactly when it matters.
How should I benchmark before buying?
Measure the metric your users feel, at the hours you actually care about, more than once. The variance between runs tells you more than any single score.
A specification sheet describes what a machine has been allocated. It does not describe how much of that allocation you will receive during your busiest hour, and on shared infrastructure that quantity varies with the behaviour of tenants you will never know about. That gap between stated capacity and available capacity is the entire subject.
Three things follow. Measure steal time properly, with interval sampling during the affected window, because the cumulative counter read once will reliably tell you that nothing is wrong. Separate the three mechanisms that all present as missing CPU — contention, a consumed burst allowance, and your own internal quotas — since only the first is a provider issue and only the second is invisible from your server. And benchmark against your own peak-hour experience rather than a quiet-host score, repeating the test across different times because the variance is the result.
Then choose the allocation model that matches the workload's shape rather than the one with the best headline number. Intermittent demand belongs on shared capacity and benefits from its economics. Sustained or latency-sensitive demand belongs on reserved capacity and benefits from its consistency. The mistake that costs money in both directions is buying for average utilisation, because the average is the one figure that never describes the moment your workload actually needs the CPU.
Accidentally deleting users, groups, or organizational ...
Bot farms are large networks of automated bots working ...
In the digital age, businesses and individuals generate...