Writing with Echo

Echo is trained to write in other people’s voices. The API serves the same model as the playground, and it is OpenAI-compatible.

Get an API key by signing in with Google on the keys page.

Calling it

import os

from openai import OpenAI

client = OpenAI(base_url="https://echo.fulcrum.inc/api/v1", api_key=os.environ["ECHO_API_KEY"], timeout=900)
reply = client.chat.completions.create(
    model="echo",
    messages=[{"role": "user", "content": "Explain GRPO in one paragraph."}],
    extra_body={"persona": "Emily Dickinson"},
)
text = reply.choices[0].message.content

The same request with curl:

curl https://echo.fulcrum.inc/api/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $ECHO_API_KEY" \
  -d '{"model": "echo", "persona": "Emily Dickinson", "messages": [{"role": "user", "content": "Explain GRPO in one paragraph."}]}'

Prompting tips

Pass the person’s name as persona and chat normally. For a conversation, send the whole history each time.

Echo does better with more detail, but giving it a bunch of slop as an outline will also make it output sloppier stuff.

Kinds of configuration:

  • A length, like “about 800 words”. Treat it as a rough target.
  • An outline or notes. Echo follows them closely: their order, claims and examples. More detail gets closer to what you want. Write them as plain notes, not in the target voice, and if the ideas are yours, keep your words.
  • Writing samples, for someone Echo doesn’t know (a friend, a small blogger, you). Paste a few of their posts above the request and pass their name as persona.
  • A draft to revise. Paste it and say what to change.
messages=[{"role": "user", "content": "Here are three of my blog posts:\n\n<post>…</post>\n\n<post>…</post>\n\n<post>…</post>\n\nWrite a post of about 1,000 words in my voice from these notes:\n\n…"}],
extra_body={"persona": "Your Name"},

Some failure modes to be proactive about

  • It invents evidence. On very open-ended prompts, because it is asked to speak and act as the target voice, Echo will sometimes make up quotes, citations, links, dates and statistics, and attach real people’s names to things they never said. You should verify things for correctness.
  • It will try to stick to your outline if you give it one, so if your outlines are slop, that will sometimes reinforce slop in its outputs. To avoid this, create outlines that are mainly trying to communicate ideas and do not have an opinionated style, or just pass in your own writing.

When calling Echo with agents, it’s often better to have the agents stay very factual in their prompts, i.e. a list of pure facts, and have Echo decide the wording and write the piece. For example:

Facts, in order:
- Jasmine is testing the app on Tuesday.
- Two schedulers will get 7 or 8 proposals each.
- Each proposal names a dollar saving.

Write a Teams message of about 100 words in my voice from these facts.

Reference

The base URL is https://echo.fulcrum.inc/api/v1, and the only model is echo. The API supports both Chat Completions and the Responses API.

persona. Every request needs it; one without it returns 400. When you set it, Echo formats the conversation the same way the playground does: it applies its own system prompt and puts the writer’s name at the top of each user message. It works best with widely published writers, or with writing samples in the message.

State. Echo is stateless, so you have to send the whole conversation on every turn, including the output items Echo returned last time. There’s no previous_response_id to chain requests with.

Reasoning. Before writing, Echo thinks at low effort unless you ask for more: set reasoning_effort in Chat Completions, or reasoning={"effort": ...} in the Responses API, to low, medium, high or max. At higher efforts a reply often takes a minute or more, so you’ll usually want to stream. Echo returns its reasoning separately from the final text: as reasoning items in the Responses API, and as reasoning_content in Chat Completions.

Streaming.

stream = client.responses.create(
    model="echo",
    input="Describe a plane taking off from Los Angeles in the 1960s.",
    extra_body={"persona": "Joan Didion"},
    stream=True,
)

for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="")

Length. Echo caps output at 20,480 tokens by default, thinking included. If you need more, you can set max_output_tokens in the Responses API (or max_tokens in Chat Completions) as high as 131,072.

Endpoints.

  • POST /api/v1/responses
  • POST /api/v1/chat/completions
  • GET /api/v1/models (which lists echo)

Any other path returns 403, and a bad or revoked key returns 401.