Jump to content

How to Tell if a Seedbox Is Oversold

From Pulsed Media Wiki


On a shared seedbox, many accounts use the same physical disks. "Oversold" is the word people reach for when one feels slow, but a slow seedbox has several possible causes, and only one of them is too many customers on too little hardware. This page explains how to tell them apart from inside your own account, on any provider, including Pulsed Media, which publishes measured disk figures for every plan and refreshes them every hour.

What "oversold" means on shared storage

The word covers two different things.

  • Space sold against space installed. A provider can sell more total storage than its disks hold, on the bet that accounts never fill at the same time. When that bet fails, the disk fills up. You see it as a full disk, not as slowness.
  • Disk work demanded against disk work available. Every disk can only do so much reading and writing per second. When the active accounts on a server ask for more than that, everyone queues. This is what customers feel as "oversold": stalled torrents and a web interface that times out.

Shared hosting only works because not every account hammers the disk at the same moment; The Economics of Shared Seedbox Hosting covers why that is fair and where it stops being fair. Two more distinctions matter before you measure anything.

Busy is not broken. A disk can be slow because healthy hardware is serving more work than it can keep up with (saturation), or because something is failing: a dying drive, a RAID array rebuilding after a drive swap (degradation). The remedies are different. Saturation is fixed by spreading the work or moving accounts. Degradation is fixed by repairing the hardware. A benchmark alone cannot tell you which one you are seeing.

Your own plan has limits. A shared plan can have its own per-account limits on memory, CPU and disk throughput. A result that stops at a steady number may be your own plan's ceiling, not your neighbours.

Why hard-drive servers give small random-I/O numbers

A 7200 rpm hard drive has to move its read head and wait for the platter to turn under it for every small random read. That takes several milliseconds each time, so one drive manages on the order of a hundred small random operations per second, somewhat more with deep request queues. A RAID array of several drives multiplies that, but the total stays small, and every account on the server draws from the same total at the same time.

A 4k random-read test on a shared hard-drive seedbox therefore prints a number that looks shocking next to an SSD, even when nothing is wrong. Seedbox traffic is a mix of longer sequential runs and scattered reads across many torrents, so test both kinds and do not judge a server on the 4k number alone.

How to measure from inside your account

No single number proves overselling. What works is several measurements, repeated at different times, read with their limits in mind.

Measurement What it shows What it cannot show
fio large-block and 4k tests How fast your account can read and write right now Whether the cause is other accounts, your plan's limit, or a hardware fault
PSI (/proc/pressure/io), the full line How much of the time every active process on the server was stuck waiting on disk Which account caused it
Load average How many processes are running or waiting, including those waiting on disk Whether the queue is CPU or disk, and whose processes they are
iowait (wa in top or vmstat) CPU time spent idle while disk requests were outstanding Disk load directly; it moves with how busy the CPUs are
Transfer speed over days What you actually get from your torrents and downloads Disk versus network versus the swarm you are in

fio, with settings that mean something

fio is a widely used Linux disk benchmark. If it is not installed and you cannot install it yourself, ask the provider. Run it on a test file in your home directory, with direct I/O so the server's memory cache does not answer for the disk, for long enough to get past short bursts:

fio --name=seqread --filename=$HOME/fio-test.tmp --size=2G --rw=read --bs=1M --direct=1 --ioengine=libaio --iodepth=8 --runtime=60 --time_based --group_reporting
fio --name=randread --filename=$HOME/fio-test.tmp --size=2G --rw=randread --bs=4k --direct=1 --ioengine=libaio --iodepth=16 --runtime=60 --time_based --group_reporting
rm $HOME/fio-test.tmp

The first line measures large-block sequential reads, the best case for hard drives. The second measures small random reads, the worst case. Real torrent traffic sits between the two. If fio reports that libaio is unavailable, drop the --ioengine option. Creating the 2 GB test file writes to the disk once and uses 2 GB of your quota until you delete it.

Keep it to read tests and run them a handful of times, not in a loop. A write test, or a benchmark repeated every few minutes, adds load to the very disk you are measuring, and on a shared server your neighbours are measuring it too.

Pressure stall information (PSI)

On Linux kernels that have it enabled, cat /proc/pressure/io prints two lines. some is the share of time at least one process was waiting on disk; it rises under any real load and is not, on its own, a sign of trouble. full is the share of time all active processes were waiting at once, which is what an overloaded disk looks like. The avg10, avg60 and avg300 fields average over the last 10, 60 and 300 seconds. The figures are for the whole server or virtual machine, not your account.

There is no published universal threshold for "too high". Look at the shape instead: a full figure that spikes and falls back is ordinary contention; one that stays high hour after hour, day and night, is a server that cannot keep up.

Load average and iowait

Load average, shown by uptime and top, counts processes that are running or waiting to run, plus those blocked waiting on disk, averaged over 1, 5 and 15 minutes. On a shared server it counts every account's processes. A high number means something is queuing; it does not say whether the queue is for CPU or disk, or whose work it is. iowait is a CPU statistic, not a disk one, and reads differently depending on how many cores the server has and how busy they are. Treat both as hints that tell you to look at PSI, never as proof.

What benchmark scripts get wrong

All-in-one benchmark scripts report the CPU cores and memory of the whole server or virtual machine. Your account's share is set by your plan's limits, not by what the script prints. A network speed test against one distant server measures the route to that server as much as your port.

Timing: peak, off-peak, and benchmark storms

Run the same tests at several times of day across a week and write down the results with the time. A seedbox that is fast at night and slow in the evening is sharing a busy disk at busy hours. One that is slow at every hour, with PSI full high around the clock, has a sustained problem.

Watch out for benchmark storms. When a server is new or being discussed on a forum, many customers run benchmarks on it in the same hour, and each of them measures everyone else's benchmark. A result taken during a storm says little about normal use.

Reading the results

What you see What it most likely means What to do
Large-block reads fine, 4k random low Normal hard-drive physics Nothing; this is how spinning disks behave
Results stop at the same round number whatever block size you try Your plan's own disk limit Check what your plan includes before blaming the server
Fine at night, slow in the evening, PSI full spiking at busy hours Ordinary contention on a shared disk Expect it on shared plans; a higher-priority or SSD plan narrows it
Slow at every hour, PSI full high all day, torrents stalling, web interface timing out Sustained overload, or a degraded array Open a ticket with your measurements
Fine for weeks, then a sudden drop A failing or rebuilding disk, or a new heavy workload on the server Open a ticket; neither is something you can fix from your account

When you contact support, send the date and time (with time zone) of each test, the exact commands, the full output, and the full line of /proc/pressure/io at the time. Say whether torrents stalled or the web interface timed out. The provider can see drive health data and the server's history over weeks, which an account cannot. Measured times make your report checkable instead of a feeling.

What a provider can publish

A customer can only see the server they are on. A provider can publish figures for every plan: how many accounts share a typical server, how much storage and how many drives it has, how long its disks take to answer, how much of the time processes wait on them, and how much of the time the disks are busy. Ask any provider whether it publishes these.

Pulsed Media publishes them on its live server performance page, refreshed every hour, with a raw-data JSON feed linked from the same page. The feed describes its own method this way:

Observed-utilization and workload figures measured from recently-reporting live servers (each server's own weekly and 60-day figures), fleet-wide and per service tier: avg_* is the average across servers, median_* the middle server. They describe real-world load and I/O experience, NOT hardware maximum-capability specifications or advertised limits: IOPS and throughput are each server's highest reading of the last 60 days (each reading a two-minute average), service time is measured I/O latency, disk utilization is the share of time disks were busy servicing I/O (iostat %util) -- not a percentage of hardware capacity. No totals, server counts, or absolute capacity are published, by design.

These are averages and medians per plan, so they describe the typical server on a plan. Your own measurements add the other half: the published figures tell you what a plan is built for, and your tests tell you how your server is doing.

At Pulsed Media

Pulsed Media seedboxes run on PMSS, its open-source server software, which sets limits per account through Linux control groups:

  • a memory limit for each account, so one account cannot take the server's memory;
  • CPU and disk priority weights, which only matter when accounts compete for the same CPU or disk at the same moment (the Seedbox and Storage Box Disk Priority page explains how disk priority works);
  • per-account disk limits: every plan has a sustained throughput ceiling sized to the plan, explained in the knowledge base article How Much Speed Does My Plan Include, and some plans also cap disk operations per second.

These controls share a busy server out fairly between the accounts on it. If your Pulsed Media seedbox is slow at every hour of the day, open a ticket with the measurements described above. Each plan's exact CPU, disk and memory numbers are on the Seedbox and Storage Box Resource Tiers page.

Frequently asked questions

Is a low 4k random fio result proof that a seedbox is oversold?

No. On a hard-drive server, small random reads are slow by physics, and every account shares the same drives. Compare large-block results across several times of day, and check PSI full, before drawing a conclusion.

Does a high load average mean the server is overloaded?

Not by itself. Load average counts every account's processes that are running or waiting, for CPU or for disk. It tells you something is queuing, not what or whose. PSI full on /proc/pressure/io is the better disk signal.

Why does a benchmark script show more CPU cores and memory than my plan includes?

The script reads the whole server or virtual machine. Your account's share is set by your plan's limits, which the script cannot see.

Can a provider prove it is not oversold?

It can publish measured figures for its plans and explain how it limits each account. Averages and medians still hide individual servers, so published figures and your own repeated measurements together are the best evidence available.

What should I send to support when my seedbox is slow?

The times of your tests with time zone, the exact commands and full output, the full line of /proc/pressure/io at those times, and what you noticed: stalled torrents, slow downloads, web interface timeouts.

See also