Storage Engine Architecture
Lioran S3 (Lioran Bastion) is an object storage server and media engine engineered in Rust. It provides single-node object storage with strict transactional metadata guarantees, bounded memory consumption, streaming I/O, and integrated media processing.
Architectural Principles
- Decoupled Planes: The metadata plane (RocksDB) is strictly segregated from the object data plane (filesystem fanout). Payload bytes are never stored inside the database.
- Bounded Memory Overhead: All ingestion and egress operations stream through fixed-size memory buffers (typically 64 KiB–4 MiB), guaranteeing that large multi-gigabyte payloads do not cause memory spikes or OOM crashes.
- Crash Consistency via Two-Phase Commit: Writes follow an atomic staging-and-promote protocol with automatic rollback on metadata commit failure.
- Resilience over Emulation: Clean RESTful semantics, strongly-typed TypeScript and CLI clients, and high-performance Rust internals without legacy AWS S3 XML/SigV4 protocol bloat.
Workspace Crate Architecture
The codebase is organized as a modular Rust Cargo workspace:
LioranBastion-Rust/
├── crates/
│ ├── bastion-common/ # Core error types, models, config, durability modes
│ ├── bastion-metadata/ # RocksDB transactional metadata engine & column families
│ ├── bastion-object/ # LocalObjectStore, staging, atomic promotions, streaming I/O
│ ├── bastion-protocol/ # REST request/response models, auth signing, token schemas
│ ├── bastion-media/ # Video transcoding, HLS/DASH packaging, thumbnail extraction
│ ├── bastion-server/ # Axum HTTP server, routing, middleware, background workers
│ └── bastion-cli/ # Native compiled admin and diagnostics binary
├── sdk/ # Official TypeScript / JavaScript client library (@liorans3/driver)
├── cli/ # Cross-platform CLI package (@liorans3/cli)
└── site/ # Documentation (Docusaurus) and Product Website (Next.js)
1. The Metadata Plane (RocksDB)
Metadata is managed by RocksDbMetadataStore inside crates/bastion-metadata. RocksDB is configured with dedicated Column Families (CFs) to isolate access patterns, optimize read-ahead caches, and avoid compaction interference across distinct domain entities.
graph TD
subgraph "RocksDB Embedded Engine"
CF_DEFAULT["default"]
CF_USERS["users (User Records)"]
CF_KEYS["access_keys (API Keys)"]
CF_BUCKETS["buckets (Bucket Metadata)"]
CF_OBJECTS["objects (Object Metadata)"]
CF_UPLOADS["uploads (Multipart In-Flight)"]
CF_JOBS["video_jobs (Transcode Queue)"]
CF_SHARES["video_shares (Signed Access)"]
CF_MANIFESTS["video_manifests (HLS/DASH)"]
CF_SYSTEM["system (Schema & Metrics)"]
end
Column Families Specification
| Column Family | Key Encoding | Value Stored | Indexing / Access Pattern |
|---|---|---|---|
users | user:{user_id} | UserRecord (JSON/Bincode) | Admin auth, Argon2id password hashes, user status |
access_keys | key:{access_key_id} | AccessKeyRecord | Programmatic API authentication (bk_... / sk_...) |
buckets | bucket:{bucket_name} | BucketRecord | Bucket configuration, creation date, quota, visibility |
objects | obj:{bucket}:{object_key} | ObjectRecord | Size, content-type, SHA-256 digest, storage path UUID, timestamps, custom metadata |
uploads | upload:{upload_id} | MultipartUploadRecord | Active multipart sessions, part indices, chunk hashes |
video_jobs | job:{job_id} | VideoJobRecord | Status (queued, processing, completed, failed), transcode parameters |
video_shares | share:{share_id} | VideoShareRecord | Temporary signed video tokens, expiry timestamps |
video_manifests | manifest:{video_id} | VideoManifestRecord | Master .m3u8 / .mpd adaptive stream manifests |
system | sys:{key} | SystemMetadata | Schema versioning, cluster ID, storage engine initialization state |
:::important Zero Payloads in RocksDB RocksDB holds only pointer paths (UUID-based identifiers) and cryptographic checksums for objects. Raw binary payloads are strictly written to the filesystem data plane. This keeps the database active working set in RAM and ensures sub-millisecond metadata lookups. :::
2. The Object Data Plane (LocalObjectStore)
Payload storage is implemented in crates/bastion-object. Objects are organized on disk using a two-character prefix directory fanout to prevent filesystem directory inode bottlenecks under millions of objects:
<storage_root>/
├── .staging/
│ ├── 019245a0-4b3e-7a11-b11c-d7e12f00a001.tmp
│ └── 019245a0-4b3e-7a11-b11c-d7e12f00a002.tmp
└── objects/
├── 4a/
│ └── 4a9f810b-b13c-4ef0-9e81-229048a87c12
├── d6/
│ └── d6617a00-9c0c-6c9a-ebf7-398d43cad6d1
└── ff/
└── ff2891bc-0a11-477b-8d19-817290123ef4
- Staging Directory (
.staging/): Active streaming uploads write to temporary files using UUIDv7 identifiers. - Objects Directory (
objects/<prefix>/<uuid>): Once fully ingested and verified, the file is promoted via an atomic filesystem rename into its permanent fanout location.
3. The Ingestion Pipeline (PUT / Atomic Write)
The atomic write lifecycle guarantees crash consistency and eliminates half-written or corrupted objects.
sequenceDiagram
autonumber
actor Client
participant Server as Axum Server
participant Staging as .staging/ (Disk)
participant DataPlane as objects/ (Disk)
participant MetaStore as RocksDB
Client->>Server: HTTP PUT /api/v1/buckets/{b}/objects/{key} (Stream)
Server->>Staging: Create UUID.tmp file
loop Stream Chunks (256 KiB)
Client->>Server: Body byte chunk
Server->>Staging: Write chunk & update incremental SHA-256
end
alt Strict Durability Mode
Server->>Staging: Execute synchronous fsync (sync_all)
else Balanced Durability Mode
Server->>Staging: Flush to OS page cache
end
Server->>DataPlane: Atomic rename (.staging/tmp -> objects/xx/uuid)
Server->>MetaStore: Write ObjectRecord to RocksDB
alt RocksDB Commit Succeeded
MetaStore-->>Server: OK (Transaction committed)
Server-->>Client: 200 OK (etag, size, SHA-256)
else RocksDB Commit Failed
MetaStore-->>Server: Error
Server->>DataPlane: Rollback: Delete physical payload file
Server-->>Client: 500 Internal Server Error
end
Ingestion Step Details
- Stream Ingestion: Incoming payload chunks from the Axum request stream flow into a 256 KiB intermediate buffer.
- On-the-Fly Hashing: A
Sha256hasher consumes bytes incrementally as they pass from network socket to disk, computing the cryptographic digest without requiring a second read pass. - Durability Enforcement:
- In
Strictmode (BASTION_DURABILITY=strict),tokio::fs::File::sync_all()issues a synchronousfsyncsyscall to ensure data is physically written to non-volatile storage media. - In
Balancedmode (BASTION_DURABILITY=balanced), writes rely on standard OS writeback caching for peak throughput.
- In
- Atomic Promotion: The completed staging file is renamed into
<storage_root>/objects/<hash_prefix>/<object_uuid>. On POSIX and NTFS systems, this rename operation is atomic within the same filesystem mount. - Metadata Transaction & Rollback: The
ObjectRecordis written to theobjectscolumn family in RocksDB. If RocksDB returns an I/O error or write failure, the physical object file is immediately deleted, preventing orphaned storage leaks.
4. The Egress Pipeline (GET & Byte-Range Slicing)
Retrieval operations support full-body streaming and precise HTTP Range: bytes=start-end requests for video scrubbing, audio seeking, and parallel download chunking.
sequenceDiagram
autonumber
actor Client
participant Server as Axum Server
participant MetaStore as RocksDB
participant Disk as objects/xx/uuid
Client->>Server: HTTP GET /api/v1/buckets/{b}/objects/{key} (Optional Range)
Server->>MetaStore: Lookup obj:{b}:{key}
MetaStore-->>Server: ObjectRecord (Path UUID, Content-Type, Size)
Server->>Disk: Open file handle
alt Byte-Range Requested (e.g. bytes=1048576-2097151)
Server->>Disk: Seek to offset via SeekFrom::Start(1048576)
Server-->>Client: HTTP 206 Partial Content (Streaming 1 MiB chunk)
else Full Body Requested
Server-->>Client: HTTP 200 OK (Streaming full body in 256 KiB chunks)
end
- Metadata Resolution: Fast index lookup in RocksDB verifies bucket existence, object key, permissions, and returns the physical storage UUID.
- Seek & Stream: The async file handle executes
tokio::io::AsyncSeekExt::seekto jump directly to the target byte offset without scanning preceding bytes. - Framed Streaming: Axum streams the requested byte window over the active TCP connection using chunked transfer encoding, keeping memory usage constant regardless of file size.
5. Multipart Upload Pipeline
For large files (multi-gigabyte to terabyte scale), MultipartUploadManager in crates/bastion-object provides resilient multi-part ingestion:
- Initiate: Client calls
POST /api/v1/multipart/init. Server allocates a uniqueUploadIdand records session state in RocksDB (uploadsCF). - Upload Parts: Client uploads individual chunks (typically 8 MiB–64 MiB) via
PUT /api/v1/multipart/{upload_id}/part?partNumber={n}concurrently across worker threads.- Each part is validated, hashed, stored as an independent staging part file, and registered with its ETag (part SHA-256).
- Complete: Client sends
POST /api/v1/multipart/{upload_id}/completewith the sorted part manifest.- The engine validates all part checksums, stitches part files sequentially into the final payload file, computes the aggregate object SHA-256, promotes the object to the data plane, and commits the object record to RocksDB.
- Abort / Cleanup: If aborted via
DELETE /api/v1/multipart/{upload_id}, all associated staging part files and metadata entries are purged.
6. Media Processing Engine
Lioran S3 includes built-in media transcoding and adaptive streaming capabilities (crates/bastion-media):
- Thumbnail Extraction: Automatically extracts representative JPEG/WebP preview images from uploaded video assets using frame-accurate seeks.
- HLS / DASH Packaging: Asynchronously segments MP4 video files into adaptive multi-bitrate HLS (
.m3u8playlists +.ts/.m4sfragments) and MPEG-DASH manifests. - Presigned Video Shares: Issues time-limited signed tokens allowing web video players (Video.js, Hls.js, Dash.js) to stream segmented media securely without exposing master storage credentials.
7. Crash Consistency & Failure Window Analysis
The table below details Lioran S3's deterministic behavior across all potential crash windows during active I/O:
| Failure Scenario | Timing of Crash | Resulting State upon Restart | Resolution Mechanism |
|---|---|---|---|
| Crash during staging write | In-flight HTTP PUT before completion | Incomplete .staging/*.tmp file on disk; no metadata entry | Daemon startup garbage collector scans .staging/ and removes unreferenced temp files older than TTL. |
| Crash during fsync | During sync_all call | Staging file may have un-flushed blocks; no metadata entry | File is never promoted. Client receives broken pipe and retries upload cleanly. |
| Crash after rename, before RocksDB commit | Physical file renamed to objects/xx/uuid, but crash occurs before metadata write | Physical payload exists on disk without corresponding RocksDB key | The object is invisible to all clients (read requests return 404). Clean-up scanner purges unindexed physical files during scheduled background maintenance. |
| Crash during RocksDB WAL write | RocksDB write-ahead log in flight | RocksDB automatically replays WAL during startup recovery | Atomicity guaranteed by RocksDB embedded WAL engine. Zero partial metadata records. |
| Crash during multipart upload | Between part uploads | Staging parts present on disk; upload session preserved in uploads CF | Client queries active parts via list_parts, skips already completed parts, and resumes remaining parts. |
Next Steps
- Review the Engineering Proof & Chaos Benchmark to see how this architecture withstood 5 live
SIGKILLhard kills under a 100.52 GiB load. - Explore the Production Durability Guide to tune disk synchronization settings for your hardware.
- Inspect the What's Missing & What's Next roadmap for upcoming distributed clustering and zero-copy transfer enhancements.