NyhederBRIEF
Ethereum & Altcoins

Why Disaster Recovery Is Core to Canton Validator Operations

RELEASE Sep 16, 2026 VIEWS 994 DESK Luke Streckenbach

On a public chain, if a validator goes down, it can simply resync with the network to catch back up. Canton, by contrast, is designed with native privacy in mind. Each node holds its own unique copy o...

On a public chain, if a validator goes down, it can simply resync with the network to catch back up. Canton, by contrast, is designed with native privacy in mind. Each node holds its own unique copy of the ledger based on which parties it hosts to enable this privacy. 

As a result, designing validator operations with resilience in mind is critical. Otherwise, node failures can result in the loss of the ledger. Node-level backup, failover, and tested recovery are vital to node operations on Canton. Institutional node operators (aka Managed Validator Operations or Node-as-a-Service) need to operate and safeguard their clients’ point of presence on the network to the same standard as is required of market-infrastructure connectivity. 

Figment’s approach prioritizes the two things that actually determine an institution’s exposure in this regard: whether the ledger is provably protected, and how fast the institution is back online when something breaks. We built our Canton validator operation around those questions rather than around uptime alone. What follows is the reasoning behind our design: what is and isn’t recoverable after a failure, why recovery speed is the metric that matters, and what a sound continuity posture looks like.

Canton’s Momentum in Tokenization

Canton has spent the past year turning institutional tokenization from a promising thesis into production reality. In July 2026, Tradeweb facilitated the first real-time on-chain U.S. Treasuries transaction on the network, with Franklin Templeton transferring a tokenized Treasury security to Virtu Financial against tokenized cash – no off-chain leg required. The Depository Trust & Clearing Corporation (DTCC) has moved from a successful collateral and margin optimization pilot to a partnership with Digital Asset to tokenize DTC-custodied U.S. Treasury securities. The institutions building on Canton are no longer experimenting at the edges: they are putting real assets and real client obligations onchain.

What draws these institutions to Canton is an architecture built for their requirements: privacy of positions and counterparties enforced at the protocol level, granular control over who can transact with whom, and interoperability across the applications running on the network. And participation in that network runs through a specific piece of infrastructure: the validator.

The Critical Role of a Validator on Canton

A validator is an institution’s point of presence on Canton. Within it, a ledger participant hosts the institution’s parties, maintains their view of the ledger, and submits transactions, while the validator application handles network protocol duties like sequencing and synchronization. Institutions generally benefit from a dedicated validator. A dedicated presence on the network gives them full control over their infrastructure, running only the applications they need and keeping the node’s full capacity behind their own activity.

Who protects the ledger?

For many of these institutions, this is the first time they have run, or even evaluated, blockchain infrastructure. Validator decisions arrive alongside a long list of others: custody integrations, application onboarding, connectivity to the Global Synchronizer, and key management among them. One consideration deserves more attention than it typically gets during evaluation: disaster recovery.

It gets less attention partly because of habits institutions carry over from traditional market infrastructure. Operational expectations formed in traditional market infrastructure hold that resilience and recoverability are the platform’s job: a participant at a clearinghouse or central securities depository doesn’t maintain its own copy of the golden record. The infrastructure provider does, and carries the business continuity obligations that come with that role. Canton deliberately inverts this model. There is no central operator holding a golden record. Each institution’s portion of the ledger lives on its own validator, and whoever runs that validator carries the responsibility for its continuity. An institution operating its own infrastructure carries that responsibility itself; one working with an infrastructure provider is delegating it, and should hold that provider to the standard it would hold itself.

This piece explains what that inversion means in practice: what is and isn’t recoverable after a failure, why recovery speed is the metric that matters, and what a sound continuity posture looks like.

What makes Canton’s architecture different

Canton is a privacy-enabled network with no shared global state. On Canton, activity is recorded as contracts between named parties – an asset and its owner, the two sides of a trade, the participants in a settlement workflow. Unlike most public blockchains, where every node holds a full copy of the chain, each Canton validator stores only the contracts its institution is a party to. If your institution isn’t involved in a transaction, your validator never sees it. No complete copy of the network’s state exists anywhere, by design.

The Global Synchronizer, operated by a decentralized collective of Super Validators, provides coordination: it sequences transactions, guarantees atomicity across otherwise separate parts of the network, and maintains the ordering that keeps every participant’s view consistent. What it is not is a global archive of participants’ private state. It can order transactions without reading them because it handles only encrypted transaction messages and the routing information needed to sequence and deliver them, but never the contents of your contracts, your positions, or your counterparty relationships.

This is Canton’s central proposition for institutions. Confidentiality of positions, counterparties, and transaction details is enforced at the architecture level. Canton’s Digital Asset Modeling Language (Daml) contracts specify exactly which parties can see or act on any piece of data, and the network’s structure ensures no one else ever holds it. (For a deeper look at Canton’s architecture, see our Canton First Look.)

The flip side of privacy: your data lives with you

The network is not your backup

On most public blockchains, every node holds the same full copy of the chain. Operating those nodes well is its own discipline, but with respect to the data itself, no single node is load-bearing. If the node an end user submits transactions through becomes unavailable, they can route to another and carry on. Even a node lost entirely can be rebuilt by replaying the chain from its peers. The network itself is the backup.

On Canton, that interchangeability doesn’t exist. An institution’s validator maintains a view of the ledger that the rest of the network cannot reproduce: there is no equivalent node to fail over to, and no full chain to replay from. Whatever redundancy protects that data is redundancy the validator operation has built itself. That said, the picture is more nuanced than “lose the node, lose everything,” and the nuance matters.

Data loss is forgiving, but downtime is not

The network does provide one meaningful assist. The Global Synchronizer retains a rolling window of transaction history (currently 30 days) and a validator restored from a recent backup automatically catches up on the transactions it missed between the backup and the point of failure. This automatic catch-up applies to Canton Coin activity processed through the Global Synchronizer, but recovery of third-party application contracts, tokenized securities and similar instruments depends on the operator’s own database backups, not the synchronizer. 

For a well-run validator, this makes outright data loss an unlikely outcome: with sound, recent backups in place, the network closes the gap on its own. However, what the network cannot give back is time. While a validator is down, its institution is dark on Canton. Transactions cannot be submitted and every workflow that runs through the node is stalled, including the legs its counterparties are waiting on. The synchronizer will hold a recovering node’s place for up to 30 days; clients, counterparties, and settlement deadlines will not. In practice, the measure of a continuity operation on Canton is not whether the ledger survives – it is how quickly the institution is back online.

The one unrecoverable scenario

There is one failure mode that is absolute rather than fast-or-slow, and it is worth explicitly calling out because it is unique to how Canton distributes data. The synchronizer’s retained history lets a node replay forward from its own last known state; it does not allow anyone to reconstruct a validator’s ledger from nothing. Canton’s validator operations documentation is explicit on this point: if no database backup, no identities backup, and no externally held keys survive a failure, the validator’s secret keys cannot be recovered. 

Unlike other validator keys, which can be rotated or held in an external key management system, the namespace root key allows for neither. If it is lost, the parties it controls cannot be recovered through any protocol mechanism. Without the namespace root key, there is no way to prove ownership of the assets they controlled. This is not a scenario any professionally operated validator should ever face. It is, however, why backup practice on Canton carries stakes it doesn’t carry elsewhere: whether those backups actually restore determines whether your institution’s assets survive.

For an institution, then, two questions define the whole topic. Is the ledger provably protected (e.g., do backups actually restore, does identity material survive the node)? And, when recovery is needed, how fast does it run? The takeaway is not that Canton is fragile, but that on Canton, the quality of your continuity posture determines whether an incident means minutes of downtime or days of disruption – and, at the extreme, whether recovery is possible at all.

What is and isn’t recoverable

Start with the first of those two questions: what survives a failure. The answer depends on what was lost and what was backed up. The distinctions are important because they drive the entire backup strategy.

Third-party assets. Tokenized securities, stablecoins, and application contracts exist as Daml contracts on your validator, and the broader network deliberately does not hold copies of them. They cannot be recovered through the network. Their continuity depends on the operator’s own database backups; beyond that, restoring lost assets would require coordination with the relevant application provider or issuer, with no guaranteed outcome.

Transaction history and audit records. Local to the validator. These are recoverable from your own backups, supplemented by the synchronizer’s automatic catch-up for the gap since the last backup within the retention window. If backups fail, this history is gone. There is no network copy to request.

Canton Coin and Canton Name Service entries. These are the exception to the “your data lives with you” rule, because the Super Validators are themselves stakeholders in Canton Coin contracts. If a validator suffers catastrophic data loss but its identities backup (the file containing the node’s namespace and signing keys) survives, a new validator can be deployed with those keys, and the Super Validators will assist recovery by providing the contracts they are party to. This restores Canton Coin balances and Canton Name Service entries for the parties the validator hosts. It does not restore transaction history, and it does not extend to contracts the Super Validators have no visibility into – which is, by design, everything else.

Keys and identities. The foundation under everything above. Keys held only on the failed node with no identities backup and no external key management system are unrecoverable, and so is everything they controlled.

What was lost Recoverable from the network? What recovery depends on
Validator node, with a recent database backup intact Yes. The gap since backup is auto-recovered via the synchronizer (30-day window) Backup freshness and a tested restore process
All node data, with only an identities backup surviving Partially. Canton Coin balances and CNS entries only The identities backup being stored securely off-node
Third-party asset contracts No The operator’s own database backups
Transaction history and audit trail Partially. Recoverable from backup and supplemented by synchronizer catch-up The operator’s own database backups
Keys (no backup, no external KMS) No. Asset ownership cannot be proven Key management and identities backup practices

Survival, though, is only the floor. What separates operators is how quickly each of these recovery paths runs when it matters, and that depends on preparation that goes well beyond backups. Before turning to what that preparation looks like, it is worth seeing the full scope of what a validator operator actually owns.

Owning a validator means owning more than uptime

Data continuity is the sharpest example of a broader point: a Canton validator operator carries a substantial operational surface beyond keeping a node online.

The network moves quickly. Canton and its application layer evolve on an active release cadence, and validators are expected to keep pace. An operator that falls behind risks losing compatibility with the network its institution depends on. Application deployments and their lifecycle sit with the operator. Transaction management, monitoring, and alerting, and key management (generation, storage, rotation, and the identities backups discussed above) sit with the operator, with direct consequences for both security and recoverability.

Node keys are not custody keys

Key management deserves a closer look, because it is easy to assume a custodian covers it. On Canton, asset-level keys and validator-level keys are different things. An institution can keep the keys that authorize movements of its assets with a custodian, held entirely outside the validator. But the validator carries key material of its own: the identity and namespace keys that establish the node’s place on the network and its right to host the institution’s parties. Those keys live with the validator operation, no custody arrangement covers them, and their survival is part of what makes post-failure recovery possible at all. Even with a custodian in place, the institution’s ledger record, and the key material protecting it, remains the validator operator’s responsibility.

None of this is unusual for teams that run production infrastructure. What is unusual, for institutions coming from traditional market structure, is that these responsibilities attach to participating in a market rather than operating one. Standing up a Canton validator means adding an operational surface your institution may never have managed before. This is exactly the operational surface a specialist infrastructure provider absorbs on an institution’s behalf. The upgrades, the monitoring, the key handling, and the continuity practice become a dedicated operations team’s full-time job, rather than one responsibility among many for a team with a different primary mandate.

What good ledger continuity looks like on Canton

Canton’s operator documentation and production operating experience point to a clear picture of what a sound posture requires. It is easiest to organize around the two questions from earlier: is the ledger provably protected, and how fast does recovery run?

Protecting the ledger

Layered backups. Canton’s operator documentation recommends backing up all database instances at least every four hours plus an identities backup held securely off-node from day one. Backup frequency should be driven by the institution’s tolerance for data loss and its recovery-time targets, not by convenience. The 30-day synchronizer window is a backstop, not a strategy: the goal is backups fresh enough that the automatically recovered gap is measured in minutes. The operational details are important here too. Canton’s documentation specifies a strict ordering between component backups, exactly the kind of subtlety that separates a checkbox backup job from a restorable one.

Privacy-preserving by construction. This is the wrinkle specific to Canton: the backups themselves contain the institution’s confidential ledger data – the very data the network’s architecture is built to protect. A continuity system that scatters unencrypted ledger copies across storage tiers has quietly undone the property the institution chose Canton for. Encryption at rest and in transit, tightly scoped access controls, deliberate data residency choices, and secure handling of identity material (which contains signing keys) are requirements, not enhancements.

Recovering fast

Tested restores. A recovery time only exists once it has been measured. Continuity procedures should be exercised against realistic failure scenarios (e.g., database corruption, loss of the validator node itself, and loss of an entire hosting region) with documented recovery steps, measured recovery times, and identified gaps. The differences between scenarios matter: a database restore, a node rebuild against an intact database, and a full cross-region recovery follow different procedures on different timelines.

Network-level readiness. The most Canton-specific determinant of recovery speed. A replacement node’s return to the network depends on more than restored data: static addressing determines whether it can inherit its predecessor’s network identity or must wait on new allow-listing coordination with Super Validator operators, and the network’s re-onboarding procedures contain steps that run considerably faster for operators who have executed them before. This is where recovery times diverge most between operators who have practiced Canton-specific re-onboarding and those who have not.

Questions to ask any operator

Whether your institution operates its own validator or works with a provider, these translate into a short list of questions worth asking:

  • What are the measured recovery times and data-loss characteristics for each failure scenario?
  • Have restores been exercised end-to-end? Against which failure scenarios, and how recently?
  • How does the recovery design handle a full regional outage?
  • How are network-level dependencies (addressing, allow-listing, re-onboarding with Super Validators) handled during a recovery?
  • What is backed up, how often, and where? Are identities and keys backed up separately and securely?
  • How does the backup and recovery system preserve confidentiality (encryption, access controls, data residency)?

How Figment approaches Canton validator continuity

We built our validator operation with ledger continuity as a first-order design requirement because on Canton, it is one.

We have executed full disaster recovery exercises against our Canton validator infrastructure. We’ve simulated database corruption, total loss of the validator node, and a complete regional outage with every scenario restored successfully, documented step by step, and measured for recovery time and data loss. Those exercises also validated the network behavior described earlier in this piece: a restored validator reconnecting to the Global Synchronizer automatically recovers the transactions it missed since its last backup. 

The design behind those results follows the posture described above. Reserved static addressing allows a replacement node to inherit its predecessor’s network identity, avoiding re-coordination with Super Validator operators during time-critical recoveries. Validator databases are backed up on a schedule designed around tight recovery-point targets, with encrypted copies maintained both locally and in independent offsite storage, and access restricted to authorized operations personnel. Recovery procedures are managed as infrastructure-as-code: peer-reviewed, repeatable, and auditable. For institutions that require sub-second recovery-point characteristics or full regional independence, standby database replicas can be provisioned across geographies.

The result is a validator operation designed to meet the business continuity standards institutions already hold their infrastructure to without compromising the privacy architecture that brought them to Canton in the first place.

If your team is evaluating Canton, or already operates on the network and wants a second look at its continuity posture, reach out to our team to talk about validator operations and disaster recovery readiness.

The post Why Disaster Recovery Is Core to Canton Validator Operations appeared first on Figment.

Source: Luke Streckenbach · www.figment.io

Discussion

Sign in to join the discussion.