Skip to main content

Serverless Inference

Deploy state-of-the-art AI models instantly with Hyperbolic’s Serverless Inference platform. Access 25+ open-source models through a simple API, with no infrastructure to manage and pricing that scales with your usage.

Upcoming Model Deprecations

The following models are being sunset and will be removed in a future update. If you are currently using any of these models, please migrate to an alternative as soon as possible.

Why Serverless Inference?

Skip the complexity of GPU management, model deployment, and infrastructure scaling. Focus on building your application while we handle the AI infrastructure.

Key Benefits

  • Instant Deployment: Start using models in seconds, not hours
  • No Infrastructure: Zero DevOps required - we handle everything
  • Pay Per Use: Only pay for the tokens/images you generate
  • OpenAI Compatible: Drop-in replacement for existing code
  • Privacy First: Zero data retention policy

Supported Model Categories

Text Generation (LLMs)

Deploy the latest language models for chat, completion, and reasoning tasks. Available Models:
  • Llama 3.1 (8B, 70B, 405B) - Meta’s latest open models
  • Qwen 2.5 (7B, 72B) - Alibaba’s multilingual models
  • Deepseek V2.5 - Efficient reasoning model
  • Hermes 3 - Fine-tuned for conversations
  • Mistral 7B - Fast and efficient
Pricing: From $0.10 per million input tokens

Image Generation

Create stunning visuals with state-of-the-art diffusion models. Available Models:
  • Stable Diffusion XL - High-quality 1024x1024 images
  • Stable Diffusion 3.5 - Latest generation
  • FLUX.1 [schnell/dev] - Ultra-fast generation (Sunset)
  • ControlNet - Guided image generation
  • Custom LoRA - Use your fine-tuned models
Pricing: From $0.0025 per image

Vision-Language Models

Process and understand images with multimodal models. Available Models:
  • Llama 3.2 Vision (11B, 90B) - Image understanding
  • Qwen2-VL (2B, 7B) - Multimodal reasoning
Pricing: From $0.15 per million tokens

Audio Generation

Generate natural-sounding speech and process audio. Available Models:
  • Melo TTS - Text-to-speech generation (Sunset)
  • Whisper - Speech-to-text transcription (coming soon)

Pricing Tiers

Note: Each source IP is capped at 600 RPM for DDoS protection. Need higher limits? Contact sales.

Developer Experience

OpenAI SDK Compatibility

Switch from OpenAI with just 2 lines of code:

REST API

Direct HTTP access for any platform:

Streaming Support

Real-time token streaming for chat applications:

Advanced Features

Function Calling

Enable models to call external tools and APIs. Supported on 18+ models including DeepSeek, Llama, Qwen, Kimi, and GPT-OSS families. See full list →
  • Structured output generation
  • Tool integration for agents
  • JSON schema validation

Custom Parameters

Fine-tune model behavior:
  • Temperature, top_p, top_k controls
  • Max tokens and stop sequences
  • Presence and frequency penalties
  • Custom system prompts

Structured Output

Get reliable JSON responses:
  • JSON mode for consistent formatting
  • Schema enforcement
  • Type validation

Batch Processing

Optimize for throughput:
  • Batch multiple requests
  • Async processing
  • Bulk pricing discounts

Use Cases

Chatbots & Assistants

Build conversational AI with streaming responses and context management.

Content Generation

Create articles, summaries, and creative writing at scale.

Code Generation

Generate, explain, and debug code across multiple languages.

Image Creation

Design assets, generate product images, and create visual content.

Data Processing

Extract insights, classify text, and analyze sentiment.

Translation

Translate content across 100+ languages with context preservation.

Getting Started

Quick Start in 3 Steps

1. Get Your API Key

Sign up at app.hyperbolic.ai and generate an API key

2. Install SDK

3. Make Your First Request

Integration Examples

LangChain Integration

Vercel AI SDK

Gradio Interface

Deploy interactive demos with one-click Hugging Face Spaces integration.

Reliability & Compliance

Infrastructure

  • 99.9% Uptime SLA for Enterprise tier
  • Global CDN for low-latency access
  • Auto-scaling to handle traffic spikes
  • Multi-region deployment

Security

  • Zero Data Retention: Your data is never stored
  • Encrypted Connections: TLS 1.3 for all API calls
  • API Key Rotation: Regular key management
  • Compliance: Pursuing SOC 2 Type II certification

Support

  • Documentation: Comprehensive guides and examples
  • Community Discord: Active developer community
  • Email Support: Pro tier and above
  • 24/7 Support: Enterprise tier

Resources

Pricing Calculator

Estimate your costs based on usage: Based on Llama 3.1 70B pricing. Actual costs vary by model. Get Your API Key →
Migration Support Moving from OpenAI, Anthropic, or another provider? Our team can help with migration strategies and code conversion. Contact us →