Today, Fulcrum is releasing Echo, the best writing model at style imitation. Echo beats frontier models at writing tasks ranging from fiction to technical explanations, despite costing less than $5K to train.

You can try it for free: contact us to scale your usage. We plan on open-sourcing the model soon.

Persona
Prompt
user
Compare with

Modern LLMs’ personalities are broadly shaped by two pressures:

  1. Perform well on agentic tasks.
  2. Follow an assistant persona. This persona is often didactic, agreeable, and has a consistently bland style.

These pressures force models into a narrow region of mind space. We think that the space of possible model character is largely unexplored: Echo is a foray into that expanse.

before (fable 5)

Strongest attack: Claim 2's mechanism statistic is mis-measured. rejection_report compares whitespace-split unit 4-grams against regex-word reference 4-grams, so the reported 46 percent sibling overlap and the 'a few percent' control are both artefacts; with consistent tokenization sibling resamples share a median 100 percent of 4-grams, 123 of 206 pairs are byte-identical …

after (echo)

Here's what we found: all three findings hold up. The strongest objection was that we screwed up the overlap measure. When you fix that, the repetition looks even stronger.

We ran temperature and sampling tests on prompts already known to loop badly. In the baseline, all 20 looped. At temperature 0.8, 16 out of 20 still looped. With top-k sampling, 18 out of 20 looped. When we just slapped a length cap on it, 20 out of 20 looped anyway. So none of those tricks really fixes the worst cases. …

From slop to Feynman.

Echo’s capabilities

Echo can do any writing task, from literary fiction to casual chatting, in whatever style you choose.

There are no good canonical metrics for writing quality, and making good metrics for writing is difficult. However, we can measure how good models are at style imitation by using standard stylometric attribution techniques. We computed a number of well-known stylometric features1 and used them to compute a style Elo2.

On this metric we are state-of-the-art at style imitation, both on authors trained on and authors not seen in training.

Figure 1. Echo has the highest style Elo, both on authors we train on, and those we didn't. These stylometric features were completely held out from the training process—we didn't even evaluate them prior to choosing our final checkpoint.2

Many people often use AI detectors, such as Pangram, as a proxy for how “slop” a model’s writing is. We think this is misguided. In our ideal world, model writing sounds beautiful, but is easily detectable as AI-generated. However, Pangram thinks Echo’s writing is human more often than frontier models.

Figure 2. Pangram judges Echo's writing as human more often than any other model we tested. All models were evaluated on a set of user prompts taken from Echo production data.

On even slightly more detailed prompts, Echo’s writing is flagged as human most of the time.

Echo’s training recipe

We train Echo on top of Kimi K3.

We believe our training gains come from elicitation: teaching the model to activate and use its suppressed knowledge of author style. This is why training Echo is so cheap: the amount of training we do is negligible relative to the amount of data the model is pre-trained on.

SFT

In our SFT stage, we seed the model’s ability to write as different personas. To do this, we use human writing from the internet.

Synthetic templates substantially increase SFT transfer

Rather than directly training on the raw posts, we prepare synthetic templates that make the data more on-policy for the original model. We prompt a frontier model with the post, and have it create an outline of the post ideas and structure. Echo is then finetuned on the original post, conditioned on the synthetic prompt.

We found this was much more effective than direct SFT:

Figure 3. Training on posts paired with a generated outline moves voice far more than training on the plain posts.

During our SFT runs, we track our voice metric, as well as the coherence of generations (judged by an LLM rubric):

Figure 4. During SFT, our voice metric rises quickly and plateaus, at a small cost in coherence. The dashed line is the human baseline.

SFT transfers surprisingly well between authors

In controlled experiments, we found a surprising degree of learning transfer between authors, i.e., training on author B helps with author A. This supports our hypothesis that the training is eliciting the model’s latent imitation ability.

We plot the expert voice score (described below) on generations prompted for the voice of author A, against SFT tokens, for data mixes with different amounts of author A’s writing.

Figure 5. Training only on persona B raises persona A's voice score almost as much as training on persona A. Lines average three seeds.

RL

Although SFT was very effective, we found that its gains plateaued, and that training on off-policy human data often incurred a cost to the model’s reasoning quality on the task. With SFT, we are also limited to training on tasks that the human authors in our data have written about. So we do a second stage of token-level RL on a voice metric we developed. We find that this teaches the model to reason properly on the task. It lets us carry our persona gains to new distributions like chat requests.

Voice expert formulation

We want a metric for how likely it is that a text on an arbitrary topic was written by author A. To do this, we fine-tune an expert \(p_a\) on A’s writing. For a text of \(T\) tokens, \(y = (y_1, \dots, y_T)\), the expert assigns probability

\[p_a(y_1, \dots, y_T) = \prod_{t=1}^{T} p_a(y_t \mid y_1, \dots, y_{t-1})\]

But the expert’s probability alone mixes together how much the text resembles the author, and how predictable the text is in general. So we compare the expert with \(p_0\), the same model before author-specific fine-tuning, and compute a likelihood ratio.

+0.3
This
+0.1
is
−0.2
the
+0.9
most
Trump expert: 38%
original model: 0.6%
TREMENDOUS
−0.3
AI
+0.1
,
+1.2
maybe
+1.7
ever
+0.2
.
\[ \underbrace{\lambda_a(y)}_{\text{voice score}} = \frac{1}{T} \sum_{t=1}^{T} \overbrace{\Big[\; \underbrace{\color{#1a6e77}{\log p_a(y_t \mid y_1, \dots, y_{t-1})}}_{\text{expert fine-tuned on author A}} \;-\; \underbrace{\color{#7a7466}{\log p_0(y_t \mid y_1, \dots, y_{t-1})}}_{\text{same model before fine-tuning}} \;\Big]}^{\text{per-token reward}} \]
Per-token rewards on a line in Trump's voice. Each bar is the bracketed term for that token, so a token the Trump expert predicts far better than the original model earns a large reward. The numbers are illustrative.

If the text sounds exceptionally like a given author, then the expert will be much better at predicting it than the original model. We use this score to get per-token rewards.

Token-level RL reinforces voice

Starting from our SFT checkpoints, we do token-level RL using the voice metric above. We train on a more diverse set of tasks, including chat requests, and writing tasks with varying amounts of detail provided. Despite only training on 8 author-specific experts, the gains of this training generalizes across the writing distribution.

Figure 6. During RL, our voice metric keeps rising while coherence holds.

Beyond the assistant persona

Labs have largely converged on a particular form of chat assistants that have been very useful, but modern post-training seems to push current models into a basin of interaction that is limited in the diversity and goals of its models. With Echo, we wanted to explore how to train capable models with different kinds of personas.3

We hope you enjoy Echo!

Try it out
or use the API

  1. We use the following four stylometric measures:

    • The cosine Delta of the text on function words (Burrows’s Delta family, Smith & Aldridge’s cosine variant).
    • Text distortion (Stamatatos 2017): character 3-grams on text where every non-function word is masked.
    • POS n-grams: universal part-of-speech bigrams and trigrams.
    • Surface features: 11 punctuation, sentence-rhythm and lexical-richness measures.

    ↩

  2. Model A “wins” against model B if the feature vector across many stylometric features is closer to the original author’s post than model B’s. ↩

  3. We were inspired by Forethought’s work on what kinds of AI systems would be differentially useful for dealing with the intelligence explosion and Gwern’s idea of “guardian angel” AIs. ↩