Skip to main content

Storage Engine Architecture

Lioran S3 (Lioran Bastion) is an object storage server and media engine engineered in Rust. It provides single-node object storage with strict transactional metadata guarantees, bounded memory consumption, streaming I/O, and integrated media processing.


Architectural Principles​

  1. Decoupled Planes: The metadata plane (RocksDB) is strictly segregated from the object data plane (filesystem fanout). Payload bytes are never stored inside the database.
  2. Bounded Memory Overhead: All ingestion and egress operations stream through fixed-size memory buffers (typically 64 KiB–4 MiB), guaranteeing that large multi-gigabyte payloads do not cause memory spikes or OOM crashes.
  3. Crash Consistency via Two-Phase Commit: Writes follow an atomic staging-and-promote protocol with automatic rollback on metadata commit failure.
  4. Resilience over Emulation: Clean RESTful semantics, strongly-typed TypeScript and CLI clients, and high-performance Rust internals without legacy AWS S3 XML/SigV4 protocol bloat.

Workspace Crate Architecture​

The codebase is organized as a modular Rust Cargo workspace:

LioranBastion-Rust/
├── crates/
│ ├── bastion-common/ # Core error types, models, config, durability modes
│ ├── bastion-metadata/ # RocksDB transactional metadata engine & column families
│ ├── bastion-object/ # LocalObjectStore, staging, atomic promotions, streaming I/O
│ ├── bastion-protocol/ # REST request/response models, auth signing, token schemas
│ ├── bastion-media/ # Video transcoding, HLS/DASH packaging, thumbnail extraction
│ ├── bastion-server/ # Axum HTTP server, routing, middleware, background workers
│ └── bastion-cli/ # Native compiled admin and diagnostics binary
├── sdk/ # Official TypeScript / JavaScript client library (@liorans3/driver)
├── cli/ # Cross-platform CLI package (@liorans3/cli)
└── site/ # Documentation (Docusaurus) and Product Website (Next.js)

1. The Metadata Plane (RocksDB)​

Metadata is managed by RocksDbMetadataStore inside crates/bastion-metadata. RocksDB is configured with dedicated Column Families (CFs) to isolate access patterns, optimize read-ahead caches, and avoid compaction interference across distinct domain entities.

graph TD
subgraph "RocksDB Embedded Engine"
CF_DEFAULT["default"]
CF_USERS["users (User Records)"]
CF_KEYS["access_keys (API Keys)"]
CF_BUCKETS["buckets (Bucket Metadata)"]
CF_OBJECTS["objects (Object Metadata)"]
CF_UPLOADS["uploads (Multipart In-Flight)"]
CF_JOBS["video_jobs (Transcode Queue)"]
CF_SHARES["video_shares (Signed Access)"]
CF_MANIFESTS["video_manifests (HLS/DASH)"]
CF_SYSTEM["system (Schema & Metrics)"]
end

Column Families Specification​

Column FamilyKey EncodingValue StoredIndexing / Access Pattern
usersuser:{user_id}UserRecord (JSON/Bincode)Admin auth, Argon2id password hashes, user status
access_keyskey:{access_key_id}AccessKeyRecordProgrammatic API authentication (bk_... / sk_...)
bucketsbucket:{bucket_name}BucketRecordBucket configuration, creation date, quota, visibility
objectsobj:{bucket}:{object_key}ObjectRecordSize, content-type, SHA-256 digest, storage path UUID, timestamps, custom metadata
uploadsupload:{upload_id}MultipartUploadRecordActive multipart sessions, part indices, chunk hashes
video_jobsjob:{job_id}VideoJobRecordStatus (queued, processing, completed, failed), transcode parameters
video_sharesshare:{share_id}VideoShareRecordTemporary signed video tokens, expiry timestamps
video_manifestsmanifest:{video_id}VideoManifestRecordMaster .m3u8 / .mpd adaptive stream manifests
systemsys:{key}SystemMetadataSchema versioning, cluster ID, storage engine initialization state

:::important Zero Payloads in RocksDB RocksDB holds only pointer paths (UUID-based identifiers) and cryptographic checksums for objects. Raw binary payloads are strictly written to the filesystem data plane. This keeps the database active working set in RAM and ensures sub-millisecond metadata lookups. :::


2. The Object Data Plane (LocalObjectStore)​

Payload storage is implemented in crates/bastion-object. Objects are organized on disk using a two-character prefix directory fanout to prevent filesystem directory inode bottlenecks under millions of objects:

<storage_root>/
├── .staging/
│ ├── 019245a0-4b3e-7a11-b11c-d7e12f00a001.tmp
│ └── 019245a0-4b3e-7a11-b11c-d7e12f00a002.tmp
└── objects/
├── 4a/
│ └── 4a9f810b-b13c-4ef0-9e81-229048a87c12
├── d6/
│ └── d6617a00-9c0c-6c9a-ebf7-398d43cad6d1
└── ff/
└── ff2891bc-0a11-477b-8d19-817290123ef4
  • Staging Directory (.staging/): Active streaming uploads write to temporary files using UUIDv7 identifiers.
  • Objects Directory (objects/<prefix>/<uuid>): Once fully ingested and verified, the file is promoted via an atomic filesystem rename into its permanent fanout location.

3. The Ingestion Pipeline (PUT / Atomic Write)​

The atomic write lifecycle guarantees crash consistency and eliminates half-written or corrupted objects.

sequenceDiagram
autonumber
actor Client
participant Server as Axum Server
participant Staging as .staging/ (Disk)
participant DataPlane as objects/ (Disk)
participant MetaStore as RocksDB

Client->>Server: HTTP PUT /api/v1/buckets/{b}/objects/{key} (Stream)
Server->>Staging: Create UUID.tmp file
loop Stream Chunks (256 KiB)
Client->>Server: Body byte chunk
Server->>Staging: Write chunk & update incremental SHA-256
end
alt Strict Durability Mode
Server->>Staging: Execute synchronous fsync (sync_all)
else Balanced Durability Mode
Server->>Staging: Flush to OS page cache
end
Server->>DataPlane: Atomic rename (.staging/tmp -> objects/xx/uuid)
Server->>MetaStore: Write ObjectRecord to RocksDB
alt RocksDB Commit Succeeded
MetaStore-->>Server: OK (Transaction committed)
Server-->>Client: 200 OK (etag, size, SHA-256)
else RocksDB Commit Failed
MetaStore-->>Server: Error
Server->>DataPlane: Rollback: Delete physical payload file
Server-->>Client: 500 Internal Server Error
end

Ingestion Step Details​

  1. Stream Ingestion: Incoming payload chunks from the Axum request stream flow into a 256 KiB intermediate buffer.
  2. On-the-Fly Hashing: A Sha256 hasher consumes bytes incrementally as they pass from network socket to disk, computing the cryptographic digest without requiring a second read pass.
  3. Durability Enforcement:
    • In Strict mode (BASTION_DURABILITY=strict), tokio::fs::File::sync_all() issues a synchronous fsync syscall to ensure data is physically written to non-volatile storage media.
    • In Balanced mode (BASTION_DURABILITY=balanced), writes rely on standard OS writeback caching for peak throughput.
  4. Atomic Promotion: The completed staging file is renamed into <storage_root>/objects/<hash_prefix>/<object_uuid>. On POSIX and NTFS systems, this rename operation is atomic within the same filesystem mount.
  5. Metadata Transaction & Rollback: The ObjectRecord is written to the objects column family in RocksDB. If RocksDB returns an I/O error or write failure, the physical object file is immediately deleted, preventing orphaned storage leaks.

4. The Egress Pipeline (GET & Byte-Range Slicing)​

Retrieval operations support full-body streaming and precise HTTP Range: bytes=start-end requests for video scrubbing, audio seeking, and parallel download chunking.

sequenceDiagram
autonumber
actor Client
participant Server as Axum Server
participant MetaStore as RocksDB
participant Disk as objects/xx/uuid

Client->>Server: HTTP GET /api/v1/buckets/{b}/objects/{key} (Optional Range)
Server->>MetaStore: Lookup obj:{b}:{key}
MetaStore-->>Server: ObjectRecord (Path UUID, Content-Type, Size)
Server->>Disk: Open file handle
alt Byte-Range Requested (e.g. bytes=1048576-2097151)
Server->>Disk: Seek to offset via SeekFrom::Start(1048576)
Server-->>Client: HTTP 206 Partial Content (Streaming 1 MiB chunk)
else Full Body Requested
Server-->>Client: HTTP 200 OK (Streaming full body in 256 KiB chunks)
end
  1. Metadata Resolution: Fast index lookup in RocksDB verifies bucket existence, object key, permissions, and returns the physical storage UUID.
  2. Seek & Stream: The async file handle executes tokio::io::AsyncSeekExt::seek to jump directly to the target byte offset without scanning preceding bytes.
  3. Framed Streaming: Axum streams the requested byte window over the active TCP connection using chunked transfer encoding, keeping memory usage constant regardless of file size.

5. Multipart Upload Pipeline​

For large files (multi-gigabyte to terabyte scale), MultipartUploadManager in crates/bastion-object provides resilient multi-part ingestion:

  1. Initiate: Client calls POST /api/v1/multipart/init. Server allocates a unique UploadId and records session state in RocksDB (uploads CF).
  2. Upload Parts: Client uploads individual chunks (typically 8 MiB–64 MiB) via PUT /api/v1/multipart/{upload_id}/part?partNumber={n} concurrently across worker threads.
    • Each part is validated, hashed, stored as an independent staging part file, and registered with its ETag (part SHA-256).
  3. Complete: Client sends POST /api/v1/multipart/{upload_id}/complete with the sorted part manifest.
    • The engine validates all part checksums, stitches part files sequentially into the final payload file, computes the aggregate object SHA-256, promotes the object to the data plane, and commits the object record to RocksDB.
  4. Abort / Cleanup: If aborted via DELETE /api/v1/multipart/{upload_id}, all associated staging part files and metadata entries are purged.

6. Media Processing Engine​

Lioran S3 includes built-in media transcoding and adaptive streaming capabilities (crates/bastion-media):

  • Thumbnail Extraction: Automatically extracts representative JPEG/WebP preview images from uploaded video assets using frame-accurate seeks.
  • HLS / DASH Packaging: Asynchronously segments MP4 video files into adaptive multi-bitrate HLS (.m3u8 playlists + .ts/.m4s fragments) and MPEG-DASH manifests.
  • Presigned Video Shares: Issues time-limited signed tokens allowing web video players (Video.js, Hls.js, Dash.js) to stream segmented media securely without exposing master storage credentials.

7. Crash Consistency & Failure Window Analysis​

The table below details Lioran S3's deterministic behavior across all potential crash windows during active I/O:

Failure ScenarioTiming of CrashResulting State upon RestartResolution Mechanism
Crash during staging writeIn-flight HTTP PUT before completionIncomplete .staging/*.tmp file on disk; no metadata entryDaemon startup garbage collector scans .staging/ and removes unreferenced temp files older than TTL.
Crash during fsyncDuring sync_all callStaging file may have un-flushed blocks; no metadata entryFile is never promoted. Client receives broken pipe and retries upload cleanly.
Crash after rename, before RocksDB commitPhysical file renamed to objects/xx/uuid, but crash occurs before metadata writePhysical payload exists on disk without corresponding RocksDB keyThe object is invisible to all clients (read requests return 404). Clean-up scanner purges unindexed physical files during scheduled background maintenance.
Crash during RocksDB WAL writeRocksDB write-ahead log in flightRocksDB automatically replays WAL during startup recoveryAtomicity guaranteed by RocksDB embedded WAL engine. Zero partial metadata records.
Crash during multipart uploadBetween part uploadsStaging parts present on disk; upload session preserved in uploads CFClient queries active parts via list_parts, skips already completed parts, and resumes remaining parts.

Next Steps​