Software Engineering & Digital Products for Global Enterprises since 2006
CMMi Level 3SOC 2ISO 27001
View all services
Staff Augmentation
Embed senior engineers in your team within weeks.
Dedicated Teams
A ring-fenced squad with PM, leads, and engineers.
Build-Operate-Transfer
We hire, run, and transfer the team to you.
Contract-to-Hire
Try the talent. Convert when you're ready.
ForceHQ
Skill testing, interviews and ranking — powered by AI.
RoboRingo
Build, deploy and monitor voice agents without code.
MailGovern
Policy, retention and compliance for enterprise email.
Vishing
Test and train staff against AI-driven voice attacks.
CyberForceHQ
Continuous, adaptive security training for every team.
IDS Load Balancer
Built for Multi Instance InDesign Server, to distribute jobs.
AutoVAPT.ai
AI agent for continuous, automated vulnerability and penetration testing.
Salesforce + InDesign Connector
Bridge Salesforce data into InDesign to design print catalogues at scale.
HumanDISC
AI-powered behavioral assessments and DISC profiling for smarter hiring.
View all solutions
Banking, Financial Services & Insurance
Cloud, digital and legacy modernisation across financial entities.
Healthcare
Clinical platforms, patient engagement, and connected medical devices.
Pharma & Life Sciences
Trial systems, regulatory data, and field-force enablement.
Professional Services & Education
Workflow automation, learning platforms, and consulting tooling.
Media & Entertainment
AI video processing, OTT platforms, and content workflows.
Technology & SaaS
Product engineering, integrations, and scale for tech companies.
Retail & eCommerce
Shopify, print catalogues, web-to-print, and order automation.
View all industries
Blog
Engineering notes, opinions, and field reports.
Case Studies
How clients shipped — outcomes, stack, lessons.
White Papers
Deep-dives on AI, talent models, and platforms.
View all resources
About Us
Who we are, our story, and what drives us.
Co-Innovation
How we partner to build new products together.
Careers
Open roles and what it's like to work here.
News
Press, announcements, and industry updates.
Leadership
The people steering MetaDesign.
Locations
Gurugram, Brisbane, Detroit and beyond.
Contact Us
Talk to sales, hiring, or partnerships.
Request TalentStart a Project
Cloud and DevOpsTechnology / AI Platform

Engineering highly scalable cloud infrastructure and a custom paywall gateway for generative AI.

Discover how we engineered a highly scalable AWS Kubernetes infrastructure and a high-throughput Paywall API Gateway to monetize custom AI/ML models securely.

AWS · EKS · Kubernetes
Client: Leading Generative AI Platform
Scalable AI Model Hosting & Paywall API Gateway

Project Overview

MetaDesign Solutions was approached by a leading AI platform to re-architect their cloud infrastructure. The client needed a robust environment to host custom AI/ML models with auto-scaling capabilities to handle unpredictable traffic spikes.

Furthermore, to successfully monetize their services, they required a custom Paywall API Gateway capable of securely managing authentication, rate limiting, and precise usage metering without adding latency to the model inference.

Cloud Infrastructure Architecture

We engineered a resilient, fault-tolerant infrastructure on AWS utilizing Amazon EKS (Kubernetes). Infrastructure as Code (IaC) was implemented using Terraform, ensuring all environments were reproducible and version-controlled.

To optimize costs, we configured Kubernetes Horizontal Pod Autoscalers (HPA) alongside Cluster Autoscaler, specifically targeting GPU node groups. This allowed the infrastructure to scale up instantly during traffic surges and scale down during quiet periods.

Custom Paywall API Gateway

We developed a high-throughput API Gateway using Node.js. To ensure sub-millisecond latency for routing and rate limiting, we integrated a Redis caching layer utilizing token bucket algorithms.

The gateway securely handles API key authentication and asynchronously flushes usage metrics (such as compute time and tokens generated) to a time-series database, integrating seamlessly with their payment processor for usage-based billing.

Have a similar challenge?

Our experts can help you build custom integrations and plugins tailored to your business workflows.

Book a free consultation

Key Challenges

01

Challenge 1

Scaling GPU-intensive workloads dynamically based on unpredictable API traffic.

02

Challenge 2

Developing a low-latency API gateway capable of complex token-based metering and billing.

03

Challenge 3

Ensuring strict security and isolation for proprietary customer models and data.

04

Challenge 4

Implementing seamless CI/CD pipelines for rapid deployment of new models.

Results & Outcomes

The new architecture transformed the client's operational capabilities. The highly scalable EKS clusters now efficiently handle over 50 million API requests daily with zero downtime. The custom Paywall API Gateway successfully monetizes the platform's usage with a processing latency of less than 50ms, enabling the client to scale their business securely and profitably.

50M+
API Requests / Day
99.99%
Uptime SLA
<50ms
Gateway Latency
100%
Automated Scaling
FAQ

Frequently Asked Questions

Common questions about this topic, answered by our engineering team.
We utilize Kubernetes (EKS) with Horizontal Pod Autoscalers and Cluster Autoscaler configured specifically for GPU node groups, ensuring capacity scales up dynamically during usage spikes and scales down to optimize costs.
The custom API Gateway was built using Node.js for high-throughput asynchronous processing, coupled with Redis for sub-millisecond latency on rate limiting, authentication, and token bucket algorithms.
Every request passing through the gateway is authenticated via JWT or API keys. Usage metrics (tokens processed, compute time) are asynchronously flushed to a time-series database and synced with Stripe for accurate usage-based billing.
Off-the-shelf solutions often lack the granular control required for AI token-based metering and dynamic routing to specific model instances based on availability and hardware requirements. A custom gateway allows for deep optimization.
Security is enforced at multiple layers: VPC isolation, private subnets for EKS nodes, IAM role-based access control (RBAC), TLS encryption in transit, and continuous vulnerability scanning in the CI/CD pipeline.
Have a similar challenge?

Let's build your success story.

A 30-minute call with a principal engineer. We'll discuss your challenges, propose architecture, and outline a roadmap.

Talk to a strategist
Have a similar project? Let's discuss your requirements.
Book a call
EmailWhatsApp