What is GitHub and Why Should You Use It?
GitHub has become one of the most popular platforms for...




Deploying a 24/7 agent takes a few minutes; keeping one alive is a different discipline entirely, because the failures that kill it are cumulative and almost none of them come from the model. They come from three slow state accumulations that no dashboard shows you by default — process memory, context, and unsupervised drift — and each one takes weeks to become visible.
This is why the first version of any always-on agent works beautifully and the third week is a mystery. Nothing changed in the code. The failure was compounding the whole time, below the threshold of anything you were measuring.
The distinction worth internalising before going further: deployment is an event, operation is a state. A deployment either succeeds or fails and you know within minutes. Operation degrades continuously, in ways that are invisible until they cross a line. Everything below concerns the second thing, and it applies equally to a chatbot daemon, a scraper, a queue worker, or an autonomous agent.
A process that runs for an hour behaves differently from one that runs for a month, and the difference is not simply thirty times longer.
Short-lived processes are forgiven for their leaks. A script that starts, does its work, and exits returns every byte it allocated to the operating system, and any inefficiency disappears with it. A persistent process never gets that reset. Every unbounded cache, every listener that was never detached, every log buffer that grew because you assumed rotation was configured, every session object retained "just in case" — all of it accumulates in the same address space with no natural ceiling.
| Failure class | Typical time to surface | First visible symptom | What actually caused it |
|---|---|---|---|
| Memory growth | Days to weeks | Process disappears with no error in its own log | Unbounded retention inside the process |
| Context growth | Hours to days | Quality degrades before anything crashes | Working set exceeds what the model can attend to |
| Unsupervised drift | Weeks | Output is plausible and wrong | No human check on high-stakes actions |
| Restart loops | Minutes | Service is marked failed and stays down | Supervision policy fighting a persistent fault |
The reason these are hard to catch is structural rather than technical. Uptime monitoring answers "is the port responding". Process supervision answers "is the process alive". Neither answers "is it doing the right thing", and by the time the answer becomes obvious the damage has usually been accumulating for days.
Memory is the most common cause of a persistent process dying, and the least well understood by the people it kills.
In managed runtimes the ceiling is explicit. V8, the engine under Node.js, divides memory into a New Space for short-lived allocations and an Old Space for objects that survive several collections. The Old Space has a limit, and when the application exceeds it the runtime raises an out-of-memory error and the process exits. For most workloads this limit is generous enough that you never encounter it, which is exactly why the failure is confusing when it arrives: the host may show free memory while the process has already hit its own internal ceiling.
Two commands make this concrete rather than theoretical.
To see the actual limit your runtime is enforcing, Node.js exposes it directly:
| Check | What it tells you | When to reach for it |
|---|---|---|
v8.getHeapStatistics().heap_size_limit |
The maximum heap the runtime will allow | Once, to confirm the ceiling is what you assumed |
process.memoryUsage() |
rss, heapTotal, heapUsed, external |
On a schedule, logged, so you can see a trend |
--max-old-space-size=4096 |
Raises the Old Space limit to 4096 MB | Only after confirming the plan has the RAM to back it |
Logging heapUsed on an interval is the single highest-value habit here, because the shape of the curve tells you what you are dealing with. A sawtooth that returns to a consistent baseline is healthy garbage collection. A curve whose baseline creeps upward every cycle is a leak, and no amount of tuning will fix it — the baseline is telling you that something is being retained that should not be.
On the operating-system side there is a second mechanism to know about. When the kernel cannot satisfy a memory request even after reclaiming caches, it invokes the out-of-memory killer, which terminates one or more processes rather than allowing the machine to become entirely unresponsive. The behaviour that surprises people most is the selection: the process the kernel kills is not necessarily the process that caused the shortage. Scores weigh memory footprint and an adjustable per-process bias, so a critical service can be selected while the actual consumer survives.
Confirming it happened takes one command against the kernel log:
dmesg -T | grep -iE 'oom|out of memory|killed process' — the classic checkjournalctl -k | grep -iE 'oom|killed process' — same messages, better timestampsA line reading Killed process 1234 (mysqld) does not mean MySQL was the problem. Treat it as the starting point of the investigation rather than its conclusion. And note that a container or cgroup with its own memory limit can trigger an out-of-memory condition inside that boundary while the host still has memory free, which is why a crash on a machine that "has plenty of RAM" is a normal result rather than a contradiction.
This one has no analogue in ordinary service operations, and it is where an always-on agent diverges most sharply from a web server.
A long-running agent's working context grows as it operates. Conversation history, tool results, retrieved documents, accumulated notes — all of it competes for a finite attention budget. Research through 2026 on long-horizon agents describes the failure modes clearly, and they compound rather than announce themselves:
The measurement literature makes the scale of this concrete. METR's time-horizon work found that the task length at which models succeed half the time has been doubling roughly every seven months, with one frontier model reaching about 110 minutes. The number that matters more for anyone running a persistent agent is the other one: the horizon at which success is reliable eighty percent of the time is roughly a quarter to a sixth of that. Capability at a given task length and dependability at that task length are two different curves, and they are far apart.
A separate multi-domain evaluation across 396 tasks and 23,392 episodes measured the decay directly: average success fell from 76.3% on short tasks to 52.1% on ultra-long tasks, a drop of 24.3 percentage points. The degradation was not uniform, either — software-engineering tasks collapsed much harder than document-processing tasks, which is what you would expect when the work involves long chains of dependent state.
The practical consequence for an always-on agent is that quality degrades before anything crashes. There is no stack trace for context drift. The agent keeps responding, keeps completing tasks, and gets quietly worse, and if your only monitoring is liveness you will not notice until someone reads the output carefully.
The third accumulation is about authority rather than resources, and it is the one most often deferred because it is a design decision rather than a configuration change.
Telemetry published by Anthropic across a very large volume of tool calls gives a useful picture of how this is handled in production today: roughly 73% of tool calls involved some form of human in the loop, while genuinely irreversible actions accounted for only 0.8%. Field surveys of practitioners point the same way, with a large majority of deployments requiring human intervention within about ten steps.
Read that as a design constraint rather than a limitation. An agent running unattended for weeks needs an explicit answer to three questions, and the answer should be written down rather than assumed:
| Question | What goes wrong with no answer | A workable default |
|---|---|---|
| Which actions are irreversible? | One bad decision compounds for days | Enumerate them; gate every one behind a confirmation |
| What is the stopping condition? | The agent runs, spends, and acts with no bound | A hard budget and a step ceiling per unit of work |
| Who notices degraded output? | Quality decays silently for weeks | A sampled review of outputs on a fixed schedule |
The supervisability of an agent is a property of its design, not of the model behind it. A well-bounded agent that asks before deleting is more useful unattended than a more capable one that does not.
The instinctive fix for all three problems is to restart the process, and on Linux the standard mechanism is a service manager that supervises it.
Used correctly, this handles the first accumulation well. Used naively, it produces a second failure that is worse than the first, because a service configured to always restart will restart a process that crashes immediately, forever. Every mainstream supervisor protects against this with a rate limit — after a configured number of failures inside a configured window, the service is marked failed and stops being restarted automatically. That is the correct behaviour, and it is also the moment an unattended system goes down and stays down.
| Setting | Purpose | What to watch for |
|---|---|---|
| Restart policy | Bring the process back after an unexpected exit | Restarting on every exit code hides real bugs |
| Restart delay | Prevent a tight loop from consuming the host | The delay also sets how long you are down each time |
| Start rate limit | Stop an unfixable service from looping forever | When tripped, the service stays down until cleared |
| Memory ceiling | Contain a leak inside a known boundary | Exceeded means terminated, so make it survivable |
That leaves the operational reality most guides omit: a scheduled restart is a stopgap, not a fix. Restarting a leaking service nightly keeps it alive and guarantees you never investigate why it leaks. It is a defensible temporary measure while you find the root cause. It is a liability as a permanent policy, because it converts a visible problem into a silent one that resurfaces at the worst possible moment — typically when the restart coincides with your busiest hour.
The same logic applies to the second accumulation. Restarting the process clears context, which restores output quality, which makes restarts look like a cure. They are masking a design question: what should this agent remember across a restart, and where should that live?
Recovery for a stateful agent is not the same problem as recovery for a stateless service, because there is state worth preserving — and the wrong place to keep it.
The general pattern is a checkpoint: the agent's state is written to durable storage at defined intervals so a new process can resume rather than begin again. Established orchestration frameworks implement this in different ways, and the difference matters when you are choosing one. Some write a checkpoint after each completed step to a database, keyed on a run identifier, which allows resuming or even replaying from an earlier point. Others treat the event history itself as the source of truth, so a crashed instance continues implicitly with no checkpoint logic in your code. Frameworks built on an in-memory store persist nothing across a process restart, which is a complete and legitimate design — as long as you know that a restart is a reset.
The failure pattern worth designing against is subtle: replaying from a checkpoint re-executes work. If the checkpoint was taken before an external call completed, resuming may repeat that call. For a read that is merely wasteful. For a payment, a message send, or a trade, it is a correctness problem, and the only durable fix is making those operations idempotent or recording their completion outside the agent's own memory.
Two rules follow, and neither is expensive:
Percentages and dashboards are secondary to a simpler principle: you cannot operate what you do not measure, and liveness is the least informative thing you can measure. A process answering on a port is not evidence that it is working.
| Signal | Detects | Alert threshold that works in practice |
|---|---|---|
| Resident memory trend | A leak before it becomes a crash | Baseline rising across several days, not a single spike |
| Restart count | An unstable service treated as normal | More than one or two restarts in a day |
| Task success rate | Silent quality decay | A drop against the agent's own earlier baseline |
| Time per completed task | Context growth before it degrades output | A sustained upward drift, not one slow outlier |
| Unreviewed output age | Drift nobody has noticed | Anything past the review interval you committed to |
Getting these into a log is more useful than getting them into a dashboard, because a log is what you will grep at the moment something breaks. A single line per interval containing memory usage, restart count, and tasks completed costs nothing and answers most retrospective questions.
The inverse also holds. Alerting on a metric nobody acts on is worse than not alerting, because it teaches the team to ignore alerts. If an alert fires and the standard response is to acknowledge it, it is noise and should be removed or re-thresholded.
Plan sizing for a persistent agent follows from the accumulations rather than from the workload's average.
Because the first failure is memory, headroom matters more than peak throughput. A plan that runs a workload at fifty percent utilisation is a plan with room for the leak to develop slowly enough to be diagnosed; a plan at ninety percent converts every small regression into an incident. The relevant question is not "does it fit today" but "what happens to this workload in four weeks if nothing about it changes".
Dedicated CPU behaves differently from shared for this kind of work, and the difference only appears over time. A long-running process needs sustained, predictable access to compute rather than the ability to burst, and background interference that would be invisible in a short benchmark accumulates into measurable drift in task duration across weeks. Deterministic latency is worth more than a better average for something you intend to leave running.
The operational overhead that people underestimate is bandwidth, though not for the reason usually cited. A persistent agent makes continuous small outbound calls — model APIs, webhooks, chat platforms, monitoring endpoints — and the volume is trivial while the connection count is not. If outbound traffic is metered and eventually capped, a service that must never pause can be interrupted by an allowance rather than by a fault. Where a plan offers a choice, prefer one whose bandwidth behaviour under sustained small requests is predictable.
None of the above requires an unusual platform, and the requirements are modest enough that a small instance is sufficient. The characteristics that matter are specific though: prebuilt images that remove the setup step, persistent storage for state that must outlive a restart, a region close enough to the services the agent talks to that latency stays low, and pricing that does not change at renewal so a long-running workload can be planned against a stable number.
One current example is the OpenClaw deployment on SurferCloud, which ships as a prebuilt image so the initial install is a selection rather than a build — relevant here mainly because every hour spent on setup is an hour not spent on the supervision layer that keeps it running. The ULightHost monthly option is a 1 vCPU / 2 GB configuration with 40 GB of storage, a dedicated IPv4 address, and 30 Mbps peak bandwidth at $6.75 per month with renewal at the same price, deployed in Japan, Singapore, or the United States. Two other configurations on the same page are first-term promotional rates that rise afterwards, at $6.90 and $4.66 per month respectively — worth knowing if the workload is intended to run for years rather than weeks, because the renewal figure is the number that shows up in every subsequent month's budget.
For an agent whose memory footprint is the primary concern, the practical detail is that 2 GB is comfortable for a single supervised process with a leak contained by a memory ceiling, and inadequate if the same host also runs a database and a vector store — those belong on a plan sized for sustained load rather than a lightweight one. And if the instance is genuinely disposable — spun up, used, discarded, and rebuilt from a checkpoint — hourly billing matches that lifecycle better than a monthly term.
The general rule transfers to any provider: choose for headroom and rental stability, not for the lowest entry price, because a persistent process is priced by what it needs in week four, not by what it needed on day one. The OpenClaw deployment page sets out the configurations and the regional options in full.
Why does my agent work for a week and then start behaving differently?
Almost always one of two cumulative causes: process memory growing past a boundary it had been staying under, or context growing past the point the model can attend to reliably. Neither produces an error. The first usually ends in a crash, the second in output that is plausible and wrong.
Is a scheduled restart a legitimate fix?
As a temporary measure while you find the root cause, yes. As a permanent policy, no — it hides the leak and postpones discovery to the moment the restart does not complete in time.
How do I tell a healthy sawtooth from a leak?
Look at the baseline. If memory returns to roughly the same resting level after each collection cycle, that is healthy. If the resting level creeps upward over days, something is being retained that is not being released, and tuning the ceiling will not correct it.
My service is marked failed and will not start. What happened?
Most likely the supervisor's start rate limit was tripped by repeated failures. That is the protection working, not a fault. Fix the underlying cause, clear the failed state, and then start it again.
Can I just raise the memory limit instead of fixing the leak?
Sometimes, and it is a reasonable stopgap if you have the headroom. It is not a fix, because a leak consumes any ceiling eventually — it only changes the date. Raising a limit on an undersized host converts a contained process exit into a system-wide out-of-memory event.
What should an agent be allowed to do without confirmation?
Anything reversible. Gate anything that cannot be undone, and make the irreversible list short enough that you can keep it accurate.
Does the agent need to run continuously at all?
Often no. If its work is genuinely periodic, a scheduled run that starts, completes, and exits avoids every accumulation described here, because a process that ends returns all of its memory. Continuous operation should be a requirement, not a default.
A 24/7 agent fails for three reasons that accumulate slowly and report misleadingly: memory grows until a boundary is crossed, context grows until quality decays, and unsupervised authority compounds into decisions nobody would have approved. None of them announce themselves, and the last two do not produce any error at all.
The work that keeps a persistent agent healthy is therefore mostly anticipation rather than troubleshooting. Log memory on an interval and read the baseline rather than the peaks. Keep state outside the process you restart, and make external side effects safe to repeat. Write down the irreversible actions and gate them. Set a supervisor that recovers from crashes but stops looping, and treat scheduled restarts as a temporary measure rather than an operating model.
Do that and the deployment question becomes the easy part, which is what it should have been. The agents that survive their third week are not the ones running the best model. They are the ones whose operators assumed from the start that something would slowly go wrong, and built the measurements that would tell them when it did.
GitHub has become one of the most popular platforms for...
6 Tbps DDoS Attack Hits Hosting Provider — Why VPS Bu...
As businesses and developers look for reliable hosting ...