Skip to content
← Insights

How to Define RTO and RPO by Process: Recovery Objectives Tied to Real-World Impact

Set RTO and RPO according to each process’s impact, dependencies, and actual recovery capacity. Learn how to agree on objectives and validate them through testing.

Business and technology teams agreeing on RTO and RPO objectives for digital processes

When a digital service is disrupted, not every process faces the same consequences or can wait the same amount of time to recover. Losing a few minutes of information is also different from losing a day of operations. That is why recovery objectives should connect business needs with technical capabilities, rather than assign one uniform target to every application.

Two measures help express those needs: RTO, which sets how long a process can remain disrupted before the impact becomes unacceptable, and RPO, which defines how much data loss, expressed as time, the organization can tolerate. These are agreed objectives, not automatic guarantees. To be useful, they must be understandable, feasible, and verifiable.

RTO and RPO answer different questions

RTO and RPO answer different questions

RTO, or recovery time objective, answers: How long can we operate without this process? It is measured from the disruption until the process returns to an agreed operational level. It is not enough for a server to start up: the service must be able to perform the functions that justify its recovery.

RPO, or recovery point objective, answers: How current must the recovered data be? If the agreed RPO is one hour, the business accepts recovering data whose state is, at most, one hour before the incident. RPO does not indicate how long restoration takes; that is covered by RTO.

The two measures complement one another, but they are not interchangeable. A system might quickly recover a copy containing data that is too old, or retain recent data but take too long to resume operations. Decisions must account for both scenarios and clarify what “recovered” means: which users, transactions, or functions need to be available.

Define objectives by process, not by application

An application may support several processes with different impacts. Conversely, a process often depends on multiple applications, data, vendors, and operational tasks. Starting with a list of systems and assigning them technical priorities can therefore obscure what the business actually needs to protect.

Begin by identifying specific processes, such as accepting orders, recording payments, or responding to requests. For each one, the business owner should describe the consequences of an interruption and of data loss. Technology teams can then map the applications and dependencies needed to support it. If the same platform serves processes with different objectives, check whether they can be recovered separately or whether they share a constraint that needs to be made explicit.

This perspective also helps uncover overlooked dependencies: identity and access, communications, integrations, databases, configuration information, staff with specialist knowledge, and external vendors. Restoring the main application does not restore the process if a critical dependency is still down.

Estimate impact over time

Impact does not always occur all at once. A brief interruption may have manageable consequences, but after a certain threshold it can lead to a growing backlog, missed commitments, or an inability to operate. The analysis should describe how harm changes over time, rather than simply labeling a system “critical.”

For a practical assessment, the team can ask:

  • What stops: the operations, channels, or decisions that depend on the process.
  • Who is affected: customers, employees, partners, or other teams.
  • What builds up: pending orders, unhandled cases, or information that is no longer being recorded.
  • When the situation becomes intolerable: the point at which impact exceeds the threshold accepted by the business.
  • What data could be lost: its importance, update frequency, and whether it can be reconstructed.

Estimates should account for relevant conditions, such as periods of high activity or operational closeouts. If there is not enough data, it is better to document the uncertainty and agree on a revisable assumption than to present an apparently precise figure without a sound basis.

Agree on objectives that can be carried out

The business owner proposes how much data loss and delay are tolerable; technology assesses the resources and procedures needed to meet those limits. Operations explains how incidents are detected, who decides to activate recovery, and which tasks need to be performed. The final decision requires discussion among these functions: setting an ambitious RTO without the resources or proven capacity to achieve it does not reduce risk.

A useful agreement defines, at a minimum, the process, RTO, RPO, recovery scope, dependencies, owners, and the evidence that will demonstrate whether the objectives were met. It also states assumptions, such as which functions will be restored first or which manual procedures can temporarily sustain activity. A manual alternative may reduce impact, but it needs owners, instructions, and clear limits; it should not be assumed to be available by default.

Comparing the business need with current capacity may reveal gaps. The answer is not always to acquire technology. It may involve simplifying dependencies, improving procedures, strengthening data capture, prioritizing an essential function, or formally accepting a different level of risk. The right choice depends on the impact and the operational options that are actually available.

Check dependencies, architecture, and operations

The agreed objective must be checked against the full recovery journey. For RTO, consider detection, decision-making, access to people and systems, restoration, verification, and the resumption of the process. If any stage is missing from the plan, the projected recovery time may be unrealistic.

For RPO, review how recoverable data is created and retained, how often it is updated, and what information falls outside that mechanism. A backup is one part of a strategy, not proof that recovery can happen within the objectives. Confirm that the data is usable and that the procedure does not depend on assumptions that have not been validated.

It is also important to identify shared dependencies. If several processes need the same identity service, network, or vendor, restoring them in parallel may compete for resources or require a specific sequence. Recording these relationships makes it possible to set restoration priorities and understand which objectives can be met under different conditions.

Test, review, and refine objectives

Testing turns objectives into evidence. An exercise can walk through the procedure from start to finish, measure how long it takes to restore the agreed functions, and verify the state of the recovered data. It is not enough to confirm that a backup exists or that an instance starts: the process owner must confirm that the result supports operations.

Exercises can begin with a review of procedures and progress to more complete technical scenarios, depending on risk and team capacity. For each test, record:

  • Which process, scenario, and dependencies were tested.
  • When the disruption began and when the agreed functions were restored.
  • Which data point was recovered and what data loss was observed.
  • Which steps failed, which assumptions did not hold, and who will address each gap.

If the result exceeds the RTO or RPO, decide whether to improve capacity, change the process, or review the objective with the business. Repeating a test without resolving its findings does not build confidence. Objectives should also be reviewed when the process, architecture, activity volume, or dependencies change.

Common mistakes and a decision template

Common mistakes and a decision template

Common mistakes include assigning the same objective to every system, confusing RTO with RPO, assuming a backup is equivalent to recovery, and setting targets without business involvement. It is also risky to measure only technical availability, ignore coordination tasks, or call a test successful without validating the data and process functions.

A simple worksheet helps keep decisions traceable. It can include these fields:

  • Process and business owner: the activity being protected and who accepts the impact.
  • Impact by duration and data loss: consequences and tolerable thresholds.
  • Agreed RTO and RPO: objectives and recovery scope.
  • Applications and dependencies: the components, teams, and vendors required.
  • Procedure and manual alternative: steps, priorities, owners, and limits.
  • Most recent test and evidence: observed results, gaps, and outstanding actions.

Defining RTO and RPO by process is not about choosing ideal numbers. It is about agreeing on limits that reflect impact, checking whether the organization can meet them, and acting on any gaps. Regular reviews keep those objectives aligned with the actual service and prevent a documented expectation from being mistaken for proven capability.

Fuentes y referencias

  1. Cloud Native GlossaryCloud Native Computing Foundation
  2. Site Reliability EngineeringGoogle