Frequently Asked Questions

General

What is Trillionir AI?

Trillionir AI is a self-hosted AI inference platform that provides OpenAI and Anthropic compatible API endpoints. You get one API key that works with every open-source model.

How does auto-routing work?

When you select "auto" as the model, our system detects your task type (coding, math, reasoning, chat) and picks the best model for that task automatically.

Is there a free tier?

Yes! The free tier gives you access to SmolLM 135M with 10 requests per minute. No credit card required.

What SDKs are supported?

Any OpenAI-compatible SDK (Python, Node.js, Go, etc.) and any Anthropic-compatible SDK. Just change the base URL to https://thetrillioniar.me/openai

Pricing

How does billing work?

We use a flat-rate subscription model. Pay monthly, get unlimited requests within your tier's rate limit. No per-token charges.

Can I upgrade or downgrade anytime?

Yes! Change your plan anytime from the dashboard. Changes take effect immediately with prorated billing.

Do you offer refunds?

We don't offer refunds for partial months, but you can cancel anytime and keep access until the end of your billing period.

Technical

Which models are available?

We support Llama 3.3, Qwen 2.5, Mistral, DeepSeek R1, and more. New models are added regularly.

Is my data private?

Yes. Each user runs in an isolated Docker container. Your prompts never touch another user's data. We don't store prompts or use them for training.

What's the max context window?

Depends on the model. Llama 3.3 supports 128K tokens, while smaller models support 4K-32K tokens.

Last updated: May 2026

Contact: [email protected]