Hide your OpenAI API key from the browser
Hide your OpenAI API key and your system prompt from the browser. This proxy keeps both on the server, puts a hard cost limit on every request, and shows you the tokens each one spent.
Your API key and your prompt stay on the server
If your page calls a model directly, your API key and your system prompt ship in the bundle, and anyone can read them. This template moves the call to a Function. The page sends the conversation, and the answer streams back.
The whole template is public. Read it before you trust it.
Or scaffold it
$ npx wawesome init --template llm-proxy
$ npx wawesome deployWhat you get
A Function called chat. It has one Hono route in
src/index.ts that takes a POST with the conversation and nothing else:
const response = await fetch(PROXY_URL, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ messages: [{ role: 'user', content: 'Where is my order?' }] }),
});
The route adds your prompt from src/system-prompt.ts and calls the model with
the official openai SDK. The answer
comes back as server-sent events: a delta event for each piece of text and a
done event at the end. A request refused before the answer starts gets a
status and a JSON error instead.
The browser can’t choose the model, the temperature or the output limit. If
a request sends one, we refuse it. Otherwise anyone with your URL could spend
your key on the most expensive model. We refuse a system message too, so no
request can change your prompt.
Four limits cap every request. They are the LIMITS constant at the top of
src/policy.ts:
| Limit | Default |
|---|---|
maxBodyBytes |
24,000 |
maxMessages |
12 |
maxPromptChars |
8,000 |
maxOutputTokens |
512 |
Each request writes one line to your logs, so you can see what it spent:
usage model=gpt-4o-mini prompt_tokens=412 completion_tokens=118 total_tokens=530 ms=1843
How many people can chat at once
An answer streams for as long as the model writes, up to two minutes. For all of that time it holds one of your plan’s invocations at once. So each plan lets this many people chat at the same moment, if nothing else in the workspace is running:
| Plan | People chatting at once |
|---|---|
| Free | 4 |
| Solo | 8 |
| Studio | 24 |
| Agency | 48 |
The next person gets a 503 until a stream ends.
Run it
npx wawesome init --template llm-proxy
It offers to log you in if you aren’t. Then it asks for a Function name, chat
by default, and an App slug. If you haven’t deployed yet, it offers to change
your workspace address, which locks at your first deploy. Then it asks for two
values:
OPENAI_API_KEY, from OpenAI Platform → Dashboard → API keys. It starts withsk-. We store it write-only, so no one can read it back.ALLOWED_ORIGINS, the sites allowed to call the proxy from a browser, likehttps://acme.com,https://www.acme.com. Leave it blank and any site can call, which is fine while you build. The Function logs a warning on every request until you set it.
It opens your App’s outbound calls to OpenAI, deploys, and prints the URL to put in your frontend.
What to change first
The prompt in src/system-prompt.ts. Ours is an example support assistant for
an online store. Replace all of it. If a new prompt goes
wrong, roll the Function back to the previous version. You don’t need to
redeploy.
Then set ALLOWED_ORIGINS to your site before you go live:
npx wawesome env set ALLOWED_ORIGINS https://your-site.com
The model is gpt-4o-mini unless you set OPENAI_MODEL. The README lists the
other settings, like per-token prices and another provider.
What it doesn’t do
ALLOWED_ORIGINS stops other websites, but not someone using curl, who can
send any Origin header. To stop that, check a session your app already
issues, next to the origin check. There’s no rate limit across requests, only
the limit on each one. If you need to limit how often one caller asks, you’ll
have to add it.
- openai
- llm
- ai
- proxy
- api-key
- typescript
Ready in about a minute
Sign in with GitHub, deploy, and get a public HTTPS endpoint.