Is your feature request related to a problem? Please describe.
Users operating large Vespa corpora sometimes need to build an entirely new,
isolated content generation from an authoritative external snapshot, such as a
data lake or warehouse export.
Common reasons include:
- Adding fields that cannot be reconstructed from documents currently stored in
Vespa.
- Changing document transformations, tokenization, schemas, or field types in a
way that is unsafe to expose as a mixed-state in-place update.
- Rebuilding or healing the complete corpus.
- Building a new application or content cluster for shadow testing, canary
rollout, migration, or rollback.
Today, the generally available approach is to feed every document through the
normal document API into an empty target cluster. This is functionally correct,
but at hundreds of millions of documents it can take days and consume substantial
CPU, network, transaction-log, indexing, flush, and fusion capacity.
The normal feed path is optimized for durable online mutations. A full bootstrap
of an isolated and initially non-serving cluster has different requirements and
may be able to avoid some of this per-document overhead.
Existing Vespa mechanisms do not appear to cover this use case completely:
- Native reindexing revisits documents already stored in Vespa and recomputes
derived fields. It cannot reconstruct fields that exist only in an external
source, and it does not build an isolated corpus generation.
- Visiting can export and re-feed documents, but visiting does not provide
snapshot isolation and the restore still uses the regular feed path.
- Node-local Proton state consists of several coordinated components, including
the document store, document metadata, disk indexes, attributes, configuration
snapshots, flushed serial numbers, and transaction-log state. Copying these
directories is not a supported, topology-independent import mechanism.
- Backup or application cloning may reproduce an existing Vespa generation, but
does not build a transformed generation from an authoritative external
snapshot.
As a result, users have to build substantial orchestration around a potentially
multi-day feed operation while also handling retries, forward changes, deletion
semantics, validation, and cutover.
Describe the solution you'd like
Provide a supported bulk-bootstrap mechanism for creating a new, initially
non-serving Vespa content generation from an immutable snapshot.
The user-facing capability is more important than prescribing a particular file
format. Possible implementations could include:
- Importing a portable Vespa-native generation artifact built offline; or
- An engine-managed bulk-load mode for an empty target that builds native state
more efficiently than the regular online feed path.
A useful interface would have the following properties:
- Immutable manifest: The import is bound to an exact Vespa version,
application/schema digest, document types, indexing configuration, source
identity, and source watermark.
- Complete query-visible state: The result covers the document store,
inverted indexes, attributes, tensors, document metadata, and delete/tombstone
semantics required to produce the same query-visible result as ordinary feed.
- Topology independence: The artifact is logical or partitioned in a way that
Vespa can safely redistribute to a target cluster whose node layout differs
from the system that produced it.
- Integrity validation: Vespa rejects incompatible, incomplete, or corrupted
input using checksums, format compatibility, schema checks, and completeness
metadata.
- Checkpoint and resume: A large import reports durable progress and can
resume after process, node, or controller failure without rebuilding successful
partitions.
- Failure isolation: An incomplete or failed generation is never exposed as
query-ready.
- Forward-write compatibility: The bootstrap establishes a base watermark
and can safely converge with mutations that occurred after that watermark.
Older snapshot state must not overwrite newer accepted writes.
- Destination completion evidence: Completion means the target has applied
and made the imported state searchable, rather than only confirming that an
external producer finished uploading data.
- Operational controls: Progress, effective rate, resource consumption,
pause/resume, throttling, and terminal errors are observable through a
supported API.
- Independent activation: Import completion should not automatically change
serving traffic. Applications can validate, shadow, canary, and activate the
new target separately.
The exact method of combining the snapshot with forward mutations could be a
separate API or follow-up feature. At minimum, the bootstrap should expose a
precise base watermark so an external indexing system can perform a safe,
versioned catch-up.
An initial, smaller version of the feature could support only:
- Empty target content clusters.
- A fixed application/schema version.
- Full-document puts and deletes.
- No topology changes during import.
- Explicit finalization before the target becomes ready.
This would still provide substantial value while leaving portable prebuilt
indexes and more advanced incremental restore for later work.
Describe alternatives you've considered
-
Regular document feed into an isolated target
This is the current safe fallback and supports arbitrary external
transformations. However, every document takes the online mutation path, and
large rebuilds may take days and create substantial indexing, transaction-log,
compaction, and network load.
-
Vespa native reindexing
Native reindexing is useful when new values can be derived from document
fields already stored in Vespa. It does not handle externally sourced fields,
incompatible whole-corpus changes, or isolated generation construction.
-
In-place field backfill
This may work for additive fields that can remain unused until coverage is
complete. It is unsafe for destructive changes, cross-field invariants, or
query behavior that cannot tolerate mixed old and new documents. It can also
contend with production serving and feed SLOs.
-
Visit/export followed by re-feed
This is useful for cloning logical documents but does not provide a
point-in-time snapshot during concurrent writes. It also retains the cost of
feeding every document through the normal destination write path.
-
Copying content-node data directories or storage volumes
Proton can load its own local persisted state during node startup, but that
state is coupled to configuration, serial numbers, bucket ownership, local
sub-databases, attributes, indexes, document metadata, and transaction-log
replay. Treating local directories as a portable public artifact would be
fragile and unsafe.
-
Backup or application/data cloning
These mechanisms are valuable for recovery or reproducing an existing
generation. They do not address rebuilding transformed documents from an
external authoritative snapshot under a new schema or indexing contract.
-
Maintaining application-specific offline index builders
Users could attempt to invoke internal Vespa indexing components themselves,
but this would couple them to unsupported storage formats and internal
invariants that may change between Vespa versions.
Additional context
The desired workflow is conceptually:
- Freeze an immutable external dataset and record its source watermark.
- Build or prepare the bulk-bootstrap artifact in parallel.
- Start or allocate an isolated target Vespa application/content cluster.
- Import the snapshot into the target using the proposed capability.
- Apply or reconcile forward changes after the snapshot watermark.
- Validate document and field parity, query behavior, freshness, capacity, and
serving latency.
- Shadow or canary the new target before switching traffic.
- Retain the previous target for rollback.
This request is not asking Vespa to:
- Read directly from a particular warehouse such as BigQuery.
- Replace application-level source-cut, validation, routing, or rollback logic.
- Expose raw content-node directories as a stable public storage format.
- Automatically activate an imported generation for serving.
- Provide a general disaster-recovery backup system as part of the same feature.
The primary success criterion would be that, for a fixed logical input and
application package, bulk bootstrap produces the same query-visible state as the
ordinary feed path while materially reducing end-to-end rebuild time and resource
cost.
Additional acceptance properties would include:
- A corrupted or schema-incompatible artifact fails before activation.
- Restarting an interrupted import resumes completed work.
- A newer forward mutation is never overwritten by older snapshot state.
- Target topology differences are either supported or explicitly rejected during
preflight.
- Import completion is based on searchable destination state.
- The target remains unavailable for normal queries until explicitly finalized.
Relevant source observations from the current Vespa implementation:
-
DocumentDB initializes coordinated local state and then replays the
transaction log:
|
LOG(debug, "DocumentDB(%s): Creating database in directory '%s'", _docTypeName.toString().c_str(), |
|
_baseDir.c_str()); |
|
|
|
_feedHandler->init(_config_store->getOldestSerialNum()); |
|
_feedHandler->setBucketDBHandler(&_subDBs.getBucketDBHandler()); |
|
saveInitialConfig(configSnapshot); |
|
resumeSaveConfig(); |
|
SerialNum configSerial = _config_store->getPrevValidSerial(_feedHandler->getPrunedSerialNum() + 1); |
|
assert(configSerial > 0); |
|
DocumentDBConfig::SP loaded_config; |
|
_config_store->loadConfig(*configSnapshot, configSerial, loaded_config); |
|
// Grab relevant parts from pending config |
|
loaded_config = DocumentDBConfigScout::scout(loaded_config, *_pendingConfigSnapshot.get()); |
|
// Ignore configs that are not relevant during replay of transaction log |
|
loaded_config = DocumentDBConfig::makeReplayConfig(loaded_config); |
|
|
|
_initConfigSnapshot = loaded_config; |
|
_initConfigSerialNum = configSerial; |
|
void DocumentDB::startTransactionLogReplay() { |
|
// This configSnapshot is only used to reuse DocumentTypeRepo |
|
// and TuneFile when loading configs during replay. |
|
DocumentDBConfig::SP configSnapshot = getActiveConfig(); |
|
IDocumentSubDB* readySubDB = _subDBs.getReadySubDB(); |
|
SerialNum oldestFlushedSerial = getOldestFlushedSerial(); |
|
SerialNum newestFlushedSerial = getNewestFlushedSerial(); |
|
(void)_state->enterReplayTransactionLogState(); |
|
_feedHandler->replayTransactionLog(readySubDB->getIndexManager()->getFlushedSerialNum(), |
|
readySubDB->getSummaryManager()->getBackingStore().lastSyncToken(), |
|
oldestFlushedSerial, newestFlushedSerial, *_config_store, |
|
_owner.shared_replay_throttler()); |
-
IndexManager::loadDiskIndex loads an internal node-local disk-index
directory, but this is only one part of complete searchable state:
|
IDiskIndex::SP IndexManager::MaintainerOperations::loadDiskIndex(const std::string& indexDir) { |
|
return std::make_shared<DiskIndexWrapper>(indexDir, _tuneFileSearch, _posting_list_cache); |
|
} |
|
|
|
IDiskIndex::SP IndexManager::MaintainerOperations::reloadDiskIndex(const IDiskIndex& oldIndex) { |
|
return std::make_shared<DiskIndexWrapper>(dynamic_cast<const DiskIndexWrapper&>(oldIndex), _tuneFileSearch); |
-
Persistence iteration explicitly does not guarantee a snapshot view:
|
* The caller shall not make any assumptions on whether or not documents |
|
* that arrive to--or are removed from--the bucket in the time between |
|
* separate invocations of iterate for the same iterator id will show up |
|
* in the results, assuming that these documents do not violate the |
|
* selection criteria. This means that there is no requirement for |
|
* maintaining a "snapshot" view of the bucket's state as it existed upon |
|
* the initial createIterator call. Neither shall the caller make any |
|
* assumptions on the ordering of the returned documents. |
We would appreciate guidance on whether an existing Vespa Cloud or self-managed
capability is intended to cover this workflow, and whether the maintainers would
prefer to approach it first as an optimized empty-cluster bulk-load mode or as a
portable native generation artifact.
Is your feature request related to a problem? Please describe.
Users operating large Vespa corpora sometimes need to build an entirely new,
isolated content generation from an authoritative external snapshot, such as a
data lake or warehouse export.
Common reasons include:
Vespa.
way that is unsafe to expose as a mixed-state in-place update.
rollout, migration, or rollback.
Today, the generally available approach is to feed every document through the
normal document API into an empty target cluster. This is functionally correct,
but at hundreds of millions of documents it can take days and consume substantial
CPU, network, transaction-log, indexing, flush, and fusion capacity.
The normal feed path is optimized for durable online mutations. A full bootstrap
of an isolated and initially non-serving cluster has different requirements and
may be able to avoid some of this per-document overhead.
Existing Vespa mechanisms do not appear to cover this use case completely:
derived fields. It cannot reconstruct fields that exist only in an external
source, and it does not build an isolated corpus generation.
snapshot isolation and the restore still uses the regular feed path.
the document store, document metadata, disk indexes, attributes, configuration
snapshots, flushed serial numbers, and transaction-log state. Copying these
directories is not a supported, topology-independent import mechanism.
does not build a transformed generation from an authoritative external
snapshot.
As a result, users have to build substantial orchestration around a potentially
multi-day feed operation while also handling retries, forward changes, deletion
semantics, validation, and cutover.
Describe the solution you'd like
Provide a supported bulk-bootstrap mechanism for creating a new, initially
non-serving Vespa content generation from an immutable snapshot.
The user-facing capability is more important than prescribing a particular file
format. Possible implementations could include:
more efficiently than the regular online feed path.
A useful interface would have the following properties:
application/schema digest, document types, indexing configuration, source
identity, and source watermark.
inverted indexes, attributes, tensors, document metadata, and delete/tombstone
semantics required to produce the same query-visible result as ordinary feed.
Vespa can safely redistribute to a target cluster whose node layout differs
from the system that produced it.
input using checksums, format compatibility, schema checks, and completeness
metadata.
resume after process, node, or controller failure without rebuilding successful
partitions.
query-ready.
and can safely converge with mutations that occurred after that watermark.
Older snapshot state must not overwrite newer accepted writes.
and made the imported state searchable, rather than only confirming that an
external producer finished uploading data.
pause/resume, throttling, and terminal errors are observable through a
supported API.
serving traffic. Applications can validate, shadow, canary, and activate the
new target separately.
The exact method of combining the snapshot with forward mutations could be a
separate API or follow-up feature. At minimum, the bootstrap should expose a
precise base watermark so an external indexing system can perform a safe,
versioned catch-up.
An initial, smaller version of the feature could support only:
This would still provide substantial value while leaving portable prebuilt
indexes and more advanced incremental restore for later work.
Describe alternatives you've considered
Regular document feed into an isolated target
This is the current safe fallback and supports arbitrary external
transformations. However, every document takes the online mutation path, and
large rebuilds may take days and create substantial indexing, transaction-log,
compaction, and network load.
Vespa native reindexing
Native reindexing is useful when new values can be derived from document
fields already stored in Vespa. It does not handle externally sourced fields,
incompatible whole-corpus changes, or isolated generation construction.
In-place field backfill
This may work for additive fields that can remain unused until coverage is
complete. It is unsafe for destructive changes, cross-field invariants, or
query behavior that cannot tolerate mixed old and new documents. It can also
contend with production serving and feed SLOs.
Visit/export followed by re-feed
This is useful for cloning logical documents but does not provide a
point-in-time snapshot during concurrent writes. It also retains the cost of
feeding every document through the normal destination write path.
Copying content-node data directories or storage volumes
Proton can load its own local persisted state during node startup, but that
state is coupled to configuration, serial numbers, bucket ownership, local
sub-databases, attributes, indexes, document metadata, and transaction-log
replay. Treating local directories as a portable public artifact would be
fragile and unsafe.
Backup or application/data cloning
These mechanisms are valuable for recovery or reproducing an existing
generation. They do not address rebuilding transformed documents from an
external authoritative snapshot under a new schema or indexing contract.
Maintaining application-specific offline index builders
Users could attempt to invoke internal Vespa indexing components themselves,
but this would couple them to unsupported storage formats and internal
invariants that may change between Vespa versions.
Additional context
The desired workflow is conceptually:
serving latency.
This request is not asking Vespa to:
The primary success criterion would be that, for a fixed logical input and
application package, bulk bootstrap produces the same query-visible state as the
ordinary feed path while materially reducing end-to-end rebuild time and resource
cost.
Additional acceptance properties would include:
preflight.
Relevant source observations from the current Vespa implementation:
DocumentDBinitializes coordinated local state and then replays thetransaction log:
vespa/searchcore/src/vespa/searchcore/proton/server/documentdb.cpp
Lines 202 to 219 in c2324d1
vespa/searchcore/src/vespa/searchcore/proton/server/documentdb.cpp
Lines 640 to 651 in c2324d1
IndexManager::loadDiskIndexloads an internal node-local disk-indexdirectory, but this is only one part of complete searchable state:
vespa/searchcore/src/vespa/searchcore/proton/index/indexmanager.cpp
Lines 53 to 58 in c2324d1
Persistence iteration explicitly does not guarantee a snapshot view:
vespa/persistence/src/vespa/persistence/spi/persistenceprovider.h
Lines 290 to 297 in c2324d1
We would appreciate guidance on whether an existing Vespa Cloud or self-managed
capability is intended to cover this workflow, and whether the maintainers would
prefer to approach it first as an optimized empty-cluster bulk-load mode or as a
portable native generation artifact.