← Templates

Hide your OpenAI API key from the browser

Hide your OpenAI API key and your system prompt from the browser. This proxy keeps both on the server, puts a hard cost limit on every request, and shows you the tokens each one spent.

Your API key and your prompt stay on the server

If your page calls a model directly, your API key and your system prompt ship in the bundle, and anyone can read them. This template moves the call to a Function. The page sends the conversation, and the answer streams back.

The whole template is public. Read it before you trust it.

Or scaffold it

$ npx wawesome init --template llm-proxy
$ npx wawesome deploy

What you get

A Function called chat. It has one Hono route in src/index.ts that takes a POST with the conversation and nothing else:

const response = await fetch(PROXY_URL, {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({ messages: [{ role: 'user', content: 'Where is my order?' }] }),
});

The route adds your prompt from src/system-prompt.ts and calls the model with the official openai SDK. The answer comes back as server-sent events: a delta event for each piece of text and a done event at the end. A request refused before the answer starts gets a status and a JSON error instead.

The browser can’t choose the model, the temperature or the output limit. If a request sends one, we refuse it. Otherwise anyone with your URL could spend your key on the most expensive model. We refuse a system message too, so no request can change your prompt.

Four limits cap every request. They are the LIMITS constant at the top of src/policy.ts:

Limit Default
maxBodyBytes 24,000
maxMessages 12
maxPromptChars 8,000
maxOutputTokens 512

Each request writes one line to your logs, so you can see what it spent:

usage model=gpt-4o-mini prompt_tokens=412 completion_tokens=118 total_tokens=530 ms=1843

How many people can chat at once

An answer streams for as long as the model writes, up to two minutes. For all of that time it holds one of your plan’s invocations at once. So each plan lets this many people chat at the same moment, if nothing else in the workspace is running:

PlanPeople chatting at once
Free4
Solo8
Studio24
Agency48

The next person gets a 503 until a stream ends.

Run it

npx wawesome init --template llm-proxy

It offers to log you in if you aren’t. Then it asks for a Function name, chat by default, and an App slug. If you haven’t deployed yet, it offers to change your workspace address, which locks at your first deploy. Then it asks for two values:

  • OPENAI_API_KEY, from OpenAI Platform → Dashboard → API keys. It starts with sk-. We store it write-only, so no one can read it back.
  • ALLOWED_ORIGINS, the sites allowed to call the proxy from a browser, like https://acme.com,https://www.acme.com. Leave it blank and any site can call, which is fine while you build. The Function logs a warning on every request until you set it.

It opens your App’s outbound calls to OpenAI, deploys, and prints the URL to put in your frontend.

What to change first

The prompt in src/system-prompt.ts. Ours is an example support assistant for an online store. Replace all of it. If a new prompt goes wrong, roll the Function back to the previous version. You don’t need to redeploy.

Then set ALLOWED_ORIGINS to your site before you go live:

npx wawesome env set ALLOWED_ORIGINS https://your-site.com

The model is gpt-4o-mini unless you set OPENAI_MODEL. The README lists the other settings, like per-token prices and another provider.

What it doesn’t do

ALLOWED_ORIGINS stops other websites, but not someone using curl, who can send any Origin header. To stop that, check a session your app already issues, next to the origin check. There’s no rate limit across requests, only the limit on each one. If you need to limit how often one caller asks, you’ll have to add it.

  • openai
  • llm
  • ai
  • proxy
  • api-key
  • typescript

Ready in about a minute

Sign in with GitHub, deploy, and get a public HTTPS endpoint.

Deploy this template