doc/development/offline_transfer.md
[!warning] Offline transfer is a work in progress. It's gated behind the
offline_transfer_exports,offline_transfer_imports, andoffline_transfer_uifeature flags, all disabled by default.
Offline transfer lets a group or project be exported to an object storage bucket and imported from that bucket later, without the source and destination GitLab instances ever talking to each other directly. This is useful when the two instances can't reach each other over the network, or when export and import need to happen at different times.
Offline transfer reuses the Direct transfer architecture: the same BulkImport, BulkImports::Entity,
and BulkImports::Tracker records, the same ETL (extract, transform, load) pipeline concern, and the same NDJSON (newline-delimited JSON) relation format. This page only
describes what's unique to offline transfer. Read Direct transfer first for the shared concepts
(terminology, pipeline design, NDJSON pipeline, idempotency, and exception handling), all of which apply here unchanged.
Direct transfer's destination instance drives the whole migration: it asks the source instance's API to export a relation, polls until it's ready, then downloads it over HTTP. Offline transfer has no source instance to ask, because export and import can happen at different times, from different networks, with nothing guaranteeing the source instance is even reachable when the import runs. Instead, export and import each run independently against object storage:
metadata.json.gz
manifest once every relation has finished.Because there's no source API to call, some direct-transfer-only steps are skipped for an offline BulkImport
(bulk_import.offline_export? is true):
BulkImports::ProcessService#import_entity
enqueues BulkImports::EntityWorker
directly instead of BulkImports::ExportRequestWorker,
since there's no source instance to ask to start an export.ProcessService skips caching a source ghost user ID, for the same reason.BulkImports::Entity#pipelines
uses a different stage list:
Import::Offline::Imports::Groups::Stage
or Import::Offline::Imports::Projects::Stage, instead of BulkImports::Groups::Stage or BulkImports::Projects::Stage.
Most pipelines are reused unmodified from direct transfer (for example LabelsPipeline, MilestonesPipeline,
BoardsPipeline, UploadsPipeline), but a few are offline-transfer-specific, under the Import::Offline::Groups::Pipelines,
Import::Offline::Projects::Pipelines, and Import::Offline::Common::Pipelines namespaces, for example
Import::Offline::Groups::Pipelines::GroupPipeline and Import::Offline::Common::Pipelines::UserContributionsPipeline.BulkImport#source_type is an enum
(gitlab or offline_export) that marks a BulkImport as an offline transfer. BulkImport#offline? is an alias
for offline_export?.Import::Offline::Export
tracks one export request, with its own state machine (created/started/finished/failed) and completion
emails. BulkImports::Export,
the same per-relation export record direct transfer uses, gets a belongs_to :offline_export so relation exports
can be grouped under one Import::Offline::Export.Import::Offline::Configuration
stores the object storage provider, bucket, encrypted credentials, the export_prefix for this export, and
entity_prefix_mapping (source full path to storage entity prefix). It's polymorphic: one row is created when an
export starts, another when an import starts, each pointing at its own bucket and credentials.flowchart TD
accTitle: Export Sidekiq job hierarchy
accDescr: ExportWorker enqueues itself and calls ExportService, which enqueues RelationExportWorker. When all relations finish, ExportWorker calls WriteMetadataService.
subgraph s1["Export"]
Import::Offline::ExportWorker -- Enqueue itself --> Import::Offline::ExportWorker
Import::Offline::ExportWorker --> BulkImports::ExportService
BulkImports::ExportService --> BulkImports::RelationExportWorker
Import::Offline::ExportWorker -- All relations finished --> Import::Offline::Exports::WriteMetadataService
end
Import::Offline::Exports::CreateService
enqueues Import::Offline::ExportWorker,
which drives Import::Offline::Exports::ProcessService:
self-relation BulkImports::Export for every descendant group or project of the
entities being exported.BulkImports::ExportService
direct transfer's source instance uses, passing offline_export_id. This enqueues
BulkImports::RelationExportWorker and, from there, the same RelationBatchExportWorker /
FinishBatchedRelationExportWorker / UserContributionsExportWorker chain documented in
Direct transfer's Sidekiq jobs execution hierarchy.BulkImports::Export::MAX_CONCURRENT_RELATION_EXPORTS
(5) relation exports run at a time.Import::Offline::Exports::WriteMetadataService writes and uploads
metadata.json.gz and marks the Import::Offline::Export finished. Otherwise, Import::Offline::ExportWorker
re-enqueues itself after 5 seconds.flowchart TD
accTitle: Import Sidekiq job hierarchy
accDescr: ScheduleImportWorker calls ScheduleImportService, which enqueues BulkImportWorker.
subgraph s1["Import"]
Import::Offline::Imports::ScheduleImportWorker --> Import::Offline::Imports::ScheduleImportService
Import::Offline::Imports::ScheduleImportService --> BulkImportWorker
end
Import::Offline::Imports::CreateService
enqueues Import::Offline::Imports::ScheduleImportWorker,
which drives Import::Offline::Imports::ScheduleImportService:
metadata.json.gz from the bucket, through
Import::Offline::Imports::MetadataFileReader.BulkImports::Entity records for the requested entities and enqueues BulkImportWorker.From here, the migration rejoins the common flow described in
Direct transfer's Sidekiq jobs execution hierarchy: BulkImportWorker,
BulkImports::ProcessService, BulkImports::EntityWorker, and BulkImports::PipelineWorker, except pipelines download
relation files from object storage instead of over HTTP (see Fog adapters and object storage),
and use the offline-specific stage lists described in How it differs from direct transfer.
Offline transfer has no GraphQL API and doesn't call the source instance's API at all. It's driven entirely through
API::OfflineTransfers, documented
in the generated REST API reference:
| Endpoint | Purpose |
|---|---|
POST /offline_exports | Starts an export. Takes the object storage configuration (provider, bucket, credentials) and a list of entities to export. Gated by the offline_transfer_exports feature flag and rate-limited. Delegates to Import::Offline::Exports::CreateService. |
GET /offline_exports and GET /offline_exports/:id | Lists or shows a user's exports, through Import::Offline::ExportsFinder. |
POST /offline_imports | Starts an import from an export already sitting in object storage. Takes the object storage configuration, the export's export_prefix, and a list of entities to import, each with a destination namespace. Gated by the offline_transfer_imports feature flag and rate-limited. Delegates to Import::Offline::Imports::CreateService. |
Direct transfer downloads relation files over HTTP from the source instance, and uploads exports through a CarrierWave
ExportUpload record. Offline transfer instead reads and writes the bucket the user configured directly, through a
small Fog wrapper built for this feature:
Import::Clients::ObjectStorage
is a provider-agnostic facade. It picks a provider-specific adapter based on Import::Offline::Configuration#provider:
aws or s3_compatible uses Adapters::Aws, gcs or gcs_application_default uses Adapters::Gcs, and gcs_hmac
uses Adapters::GcsHmac. All three wrap a plain Fog::Storage.new(provider: ..., **credentials) client.
s3_compatible (for example MinIO) is only offered when the allow_s3_compatible_storage_for_offline_transfer
application setting is enabled.gcs_application_default (GCS Application Default Credentials) resolves to the service account of the instance
running GitLab, so it's never offered on GitLab.com and only offered on self-managed when an administrator has
enabled the allow_application_default_credentials_for_offline_transfer application setting.Uploading calls directory.files.create, using multipart upload above a 100 MB threshold. Downloading streams the
file in chunks through directory.files.get.
Import::Offline::ObjectKeyBuilder
is the single source of truth for where a file lives in the bucket:
<export_prefix>/<entity_prefix>/<relation>.<extension> # unbatched
<export_prefix>/<entity_prefix>/<relation>/batch_<n>.<extension> # batched
<export_prefix>/metadata.json.gz # metadata
For example: 2026-04-16_19-39-00_export_dJtnb3CV/project_1/repository.tar.gz. On export, entity_prefix is derived
directly from the portable being exported (project_1, group_5). On import, it's looked up from
Import::Offline::Configuration#entity_prefix_mapping, which is populated from the entity mapping written into
metadata.json.gz during export.
Uploading and downloading each have a dedicated strategy class, parallel to the ones direct transfer uses for HTTP:
Import::Offline::ExportUploadable
is mixed into the export services that need it (for example FileExportService, UploadsExportService). When
offline_export_id is present, it skips creating a CarrierWave ExportUpload and uploads the file straight to
object storage instead.BulkImports::FileDownloadService.for_context
is the single branch point on the import side: for an offline pipeline context, it builds
Import::Offline::Imports::ObjectStorageFileDownloadStrategy; otherwise it builds Import::BulkImports::HttpFileDownloadStrategy,
direct transfer's HTTP strategy. ObjectStorageFileDownloadStrategy streams the file from object storage,
validates its Gzip header, and enforces the bulk_import_max_download_file_size application setting the same
way the HTTP strategy does.