Get $100,000 in storage credits for your Gen-AI startup

Apply for the Grant

S3-compatible data storage
for generative AI and machine learning

Store training datasets, embeddings, model checkpoints and generated outputs with S3-compatible object storage built for Gen-AI and ML pipelines.

  • S3 — native API compatibility
  • $0 — API request fees
  • 1,462 MB/s — upload throughput (MinIO Warp)
  • $20 — starter credit, no credit card required

Built for teams training models and generating media at scale — LLMs, diffusion models, and traditional ML pipelines alike.

Built for Gen-AI and ML workflows

What you can store with Rabata

Store training data, embeddings, model artifacts and generated outputs in S3-compatible object storage built for Gen-AI and ML pipelines.

Training datasets

Store massive training corpora at any scale, from fine-tuning sets to foundation-model-scale corpora. No capacity planning or upfront provisioning required.

Training corpora stored in object storage

Model training pipelines

Feed training jobs and distributed pipelines directly via the S3 API. PyTorch, TensorFlow and Hugging Face support S3 natively.

Distributed training pipeline reading from S3

Embeddings and feature data

Store embedding vectors, feature matrices and index shards for RAG and retrieval pipelines. Processes up to 696 objects/sec in benchmark tests — metadata operations stay fast.

Embedding vectors and index shards in object storage

Generated outputs

Store the video, image and audio your models generate — not just what they train on. Same S3 endpoint, same predictable rate, no separate media pipeline to provision.

Generated video, image and audio stored in one bucket

Generating video with your model?

If you need to transcode and deliver generated video at scale, Rabata Video handles the full pipeline on the same infrastructure: one API to upload, transcode and stream.

Why teams switch

Three problems Gen-AI and ML teams hit first

AI workloads generate massive datasets — and, increasingly, massive volumes of generated output — that traditional storage systems struggle to manage.

Massive datasets

Training data routinely reaches terabytes or petabytes. Traditional file systems and NAS don't scale without hardware procurement — a config change becomes a procurement cycle

Pipeline bottlenecks

Distributed training clusters need high-throughput storage to keep GPUs continuously fed. Slow storage means idle compute — the bottleneck is storage, not compute

Infrastructure complexity

Managing distributed storage for AI adds operational overhead. Teams end up maintaining storage systems instead of building models

Hyperscaler object storage often works technically, but request fees and unpredictable egress make large-scale AI storage harder to plan around. Rabata removes API request charges entirely and bills storage and egress through one simple, predictable rate — $6/TB per month.

High-throughput storage for AI data pipelines

Storage performance directly affects GPU utilization — slow ingestion means idle compute. Benchmarked using MinIO Warp v1.0.7 against major S3 providers under identical conditions.

Metric
Rabata
AWS S3
DigitalOcean
vs AWS
Upload (PUT) 1,462 MB/s ★ 1,444 MB/s 1,440 MB/s Highest in this benchmark
Mixed workload 346 MB/s ★ 151 MB/s 179 MB/s 2.3× AWS
Small objects 696 obj/s ★ 319 obj/s 328 obj/s 2× AWS

Download throughput (GET) depends on workload pattern and varies across test configurations. Upload and mixed workload metrics reflect sustained sequential patterns most relevant to ML training pipelines.

MinIO Warp v1.0.7 · Debian 13 VM · 16 GB RAM · US-East-1 · 8 concurrent threads · Full methodology: See Full Performance Benchmark →

Connect Rabata to any Gen-AI or ML workflow in three steps

No code changes. No new agents. No proprietary connectors.

  1. Store AI datasetsUpload training data and preprocessing outputs to a Rabata bucket in eu-west-2 or us-east-1. No credit card required, $20 in starter credit included.

  2. Feed ML pipelinesTraining pipelines read data directly via the S3 API. Configure your endpoint, access key and secret key. PyTorch, TensorFlow, Hugging Face, MLflow and Spark all work out of the box.

  3. Store model artifacts and generated outputsWrite checkpoints, embeddings, and generated video/image/audio back via the same S3 API. Retrieve any artifact instantly — no rehydration, no queues.

See how much you'd save

One tier. No API request fees. $20 starter credit, no credit card required.

S3 Storage — $6/TB One price for storage and chargeable egress. Incoming traffic is free. Each billing month includes egress up to 3× your average monthly storage; only traffic above that allowance is charged. The allowance resets monthly and does not roll over.
100 TB
250 TB
Included egress (3× average storage) 300 TB
Chargeable egress above allowance 0 TB
Storage cost $600
Egress cost $0
Estimated monthly total $600 / month
Rabata $0 / year
Microsoft Azure $0 / year
Amazon S3 $0 / year
Google Cloud $0 / year

Get up to $100,000 in storage credits

If you're building Gen-AI or machine learning products, apply for Rabata's startup grant to cover your storage costs while you scale.

Storage costs shouldn't slow down your Gen-AI product. The Rabata grant gives early-stage teams up to $100k in credits to build and scale without infrastructure cost becoming a bottleneck.

  • Up to $100,000 in storage credits
  • No lock-in or long-term commitment
  • Works with your existing S3-compatible stack

Works with popular AI and data tools

Any tool with S3 support works with Rabata. Configure endpoint, access key and secret key — nothing else needed.

  • PyTorch — S3 dataset loaders (s3://bucket/path) with torch.utils.data and torchdata S3 connectors. No additional setup.
  • TensorFlow / Keras — tf.io and tf.data pipelines support S3 via the filesystem plugin. Works with TFDS and custom training pipelines.
  • Hugging Face Datasets — load and stream datasets using storage_options with your Rabata endpoint. No code changes beyond endpoint config.
  • MLflow — store experiment artifacts, models and metrics using the S3 artifact store URI in mlflow.set_tracking_uri.
  • Apache Spark — configure the S3A connector with your Rabata endpoint and credentials. Supports large-scale distributed dataset processing.
  • Any S3 tool — if your stack supports S3, it works with Rabata. Change the endpoint — nothing else.

Moving an existing pipeline over takes under 30 minutes — step-by-step migration guide. See full pricing and estimate your cost.

Secure AI data storage

  • IAM access control — role-based policies control which users, services and training jobs can access specific buckets. Separate production datasets from experiment storage.
  • Private buckets by default — all buckets are private unless explicitly configured otherwise. Training data, model artifacts and generated outputs are inaccessible to unauthorized requests.
  • Versioning & immutability — protect training datasets and model checkpoints from accidental overwrite. Roll back to a prior dataset version or checkpoint to reproduce a specific training run.
  • SigV4 + HTTPS · regional data placement — all API requests use AWS Signature Version 4. All connections encrypted via HTTPS. Buckets available in eu-west-2 (EU) and us-east-1 (US) — supports GDPR data residency requirements for European teams.

Frequently asked questions about AI data storage

AI data storage is the infrastructure used to store, access and manage the large datasets required for machine learning and generative AI — raw corpora, preprocessed splits, embedding vectors, model checkpoints, generated outputs and experiment artifacts. Object storage is the standard approach because it scales without hardware limits, supports high-throughput parallel access and integrates natively with ML frameworks via the S3 API.

Storage is billed at $6/TB per month, with egress included up to 3× your average monthly storage each month — covering most training pipelines that re-read data across epochs. Only usage above that allowance is billed, at the same rate. Full breakdown in the billing documentation.

Yes. Configure your Rabata endpoint (s3.eu-west-2.rabata.io or s3.us-east-1.rabata.io) with your access key and secret key — no SDK changes, no custom connectors. Works with PyTorch's torchdata S3 connector, TensorFlow's tf.io filesystem plugin and Hugging Face's storage_options parameter.

Yes. Versioning keeps prior states of a dataset or checkpoint recoverable, and Object Lock immutability protects them from accidental overwrite — so you can trace which dataset version or checkpoint produced a specific training run.

Under 30 minutes. Sign up (no credit card required, $20 in starter credit included), generate access keys, create a bucket and update your S3 endpoint config. No new agents, no code changes.

The best object storage for AI workloads combines high upload throughput, low-latency parallel reads, S3 API compatibility and predictable pricing. Rabata delivers 1,462 MB/s upload throughput and 346 MB/s on mixed workloads — 2.3× faster than AWS S3 in the same benchmark — with no API request fees.

ML pipelines read training data from Rabata via the S3 API during each epoch. PyTorch, TensorFlow, Hugging Face Datasets and MLflow all support S3 natively — configure your access key, secret key and regional endpoint. Repeated reads across epochs are covered by the included egress allowance for most training patterns.

Yes. Store raw datasets, preprocessed splits, feature stores, embedding indexes, model artifacts and generated outputs in a single bucket hierarchy. Access data from any point in your pipeline — preprocessing with Spark, training with PyTorch or TensorFlow, experiment tracking with MLflow — all via the same S3 API endpoint and the same storage rate.

Rabata S3 storage holds whatever your models produce — including generated video, image and audio. If you need to transcode and deliver that output at scale, Rabata Video runs on the same infrastructure and handles the full pipeline.