All work

Case studyPKNIC2026

A Bilingual LLM Eval Gate for Content Generation

I cut posting one event to three channels from 30 minutes to under 5. One brief becomes a post for each channel, in English or Korean, previewed the way the app will show it, and my feedback shapes the next draft.

Role
Solo: product, eval design, model selection and the build
Result
  • 3 channels in under 5 minutes, down from 30 or more
  • 0 invented dates, prices or venues in the final Korean run
Compared
3 cloud and 7 local models head to head; 2 more ruled out early
Model
HyperCLOVA X 1.5B, 1.1GB, running locally

01The problem

Four channels, and the same fixes every time.

PKNIC, the community I run events for, posts on four channels: LinkedIn, Instagram, Circle and KakaoTalk, Korea's messenger app. Each has its own audience, voice and length.

For every event I wrote a draft, asked an LLM for a version per channel and reviewed each one. The writing often did not fit the channel's format, and nothing learned from my corrections, so I made the same fixes every time. Three channels took me at least 30 minutes to write, review and check on the platform.

So I built a tool that turns one brief into a post for each channel, shows each post the way its app will, and learns from my feedback. The tool itself was simple. Choosing the model was the hard part.

02Choosing the model

The cheapest model that writes well, in two languages, on 8GB.

My priority was keeping token costs as low as possible.

Round 1: cloud models

I started with hosted models from Anthropic, OpenAI and Google. They understood the context well. But many posts are simple, and I did not want to pay for tokens on them, even on the cheaper tiers. They also invented facts. The marks below are mine, on their actual output.

LinkedIn · one brief, three cloud models · Jul 24, 2026

Coach Park mentoring session with sponsor KOTRA July 26th 2026, 12PM PST, Seattle Convention Center Audience: University Students, Recent grads

Claude Haiku 4.5 1,531 chars

I've sat across from hundreds of hiring managersinvented first-person credential, and here's what they won't tell you in a career fair… 📍 Seattle Convention Center 📅 SaturdayJul 26, 2026 is a Sunday, July 26th, 2026

Invented a weekday, and got it wrong. Also 12 ** markdown artifacts that render as literal asterisks.

Gemini 2.5 Flash 1,596 chars

Imagine gaining direct insights and mentorship from Coach Park, a proven leader whose strategies have propelled countless individuals and organizations forwardnot in the brief. Coach Park's mentoring sessions are renowned for their practical adviceinvented reputation…

Invented a reputation for a real person. 14 ** artifacts, and nearly twice LinkedIn's comfortable length.

Claude Sonnet 4.6 501 chars

박 코치 멘토링 세션 — KOTRA 후원, 시애틀에서 열립니다.English brief; the others wrote English KOTRA는 한국 무역·투자 진흥을 이끄는 정부기관입니다.world knowledge, not the brief 📅 2026년 7월 26일 (일)correct weekday | 오후 12시 PST

The best writing of the three, in Korean for an English brief, with a fact from outside it.

A stronger model was not a safer one: the weekday Haiku invented was wrong, and the fact Sonnet added happened to be right, which a reviewer cannot tell apart without checking.

Round 2: local models on an 8GB laptop

Local models such as Qwen and Llama cost nothing to run, but they run on my MacBook Air M3, which has 8GB of memory. Models above about 4B parameters barely ran, so I needed models that are small and still write well.

Then Korean broke them. On Korean briefs, small general models answered in Thai, Chinese or even Russian, and Llama models mixed scripts mid-word. My community posts in Korean too, so I added four models trained on Korean data and compared all seven on one bilingual brief across six channels. An automatic rubric scored each post for length, format and tone, and I counted the rules each model broke.

ModelOriginMean rubricRule violationsWhat stood out
qwen2.5:3bGeneral90.06Top rubric score, most violations
hyperclovax:1.5bKorean, Naver89.54Chosen. Smallest at 1.1GB, 3.9 to 5.1 seconds a post against 9 to 59 for the rest, and no script corruption
gemma2:2bGeneral87.74Invented “14일 (금)”; the date is a Monday
kanana:2.1bKorean, Kakao86.83Fewest violations
exaone3.5:2.4bKorean, LG85.04Romanized the speaker's name into a different name
llama3.2:3bGeneral78.74Script corruption: “세attle 대학교”, “기회가 arriving!”
bllossom:3bKorean fine-tune of Llama74.44, plus a timeoutInherited Llama's corruption, and invented links like t.co/shortened_url
7 local models across 6 channels, Sep 21, 2026. Earlier, gemma3:4b worked but was slow on 8GB, and qwen3:4b, a reasoning model, took 357 seconds for one post.

I chose HyperCLOVA X, a Korean model from Naver: the smallest and fastest of the seven, with no mixed scripts. But no model, Korean or general, stopped inventing facts. Every one broke 3 to 6 rules its prompt stated, and the rubric could not see it: it scored a post with [Insert Price] 98 of 100.

03What I decided

Three rules for working with the model.

  1. 1

    Facts come from form fields

    The date, time, venue, price, speaker, event type and link come from fields I fill in, and code writes them the same way every time. Voice, tone and structure stay with the model, tailored to each channel.

  2. 2

    Feedback beats a perfect prompt

    No instruction is perfect. I tag what is wrong with each draft, such as its tone, its length or a wrong fact. Voice tags shape the model's next drafts, and fact tags go to a list of defects to fix in code.

  3. 3

    Say what to do

    Telling the model what to do worked better than a list of what never to do.

04How it works

One brief, a preview for each channel, then my feedback.

I fill in the event details and pick the channels. The model writes each post without seeing the facts. Code deletes any sentence that claims a fact the form did not give, lays the facts out for each channel and shows the preview. I approve the post, or tag what is wrong.

my feedback tags shape the next draftsCODEMODELCODECODEMEForm fieldsdate, time, venue, price,speaker, type, linkWrites the postvoice, tone, structurefor each channelSentence scrubdeletes unbacked facts,flags every deletionLayout and previewfacts block, shownas each app shows itReview and tagapprove, or tag whatis wrong and whythe facts skip the model and go straight to the layout

The scrub, on sentences that actually shipped

Every sentence below came from a real run. The facts belong only in the details block, which code writes, so deleting a whole sentence costs the reader nothing.

  • The event will take place at our state-of-the-art facility in Seoul, South Korea.Deleted: the event is in Seattle
  • Register now and get a free resource book as a thank you gift.Deleted: no price was given
  • Save your seat on LinkedIn today.Kept: about the reader, not the event

If a model's opening line is deleted, code writes one from the form.

A layout for each channel's reader

Correct is not the same as readable. I read each output as someone scrolling that app would, and gave each channel its own layout in code: one “When” line for the date and time, topics grouped under one heading, blank lines between blocks, ▶ markers in KakaoTalk's group-notice style, and “link in bio” on Instagram, whose captions don't make links clickable.

KakaoTalk read in a group chat on a phone

Before: Don't miss out on this great opportunity to learn from the best! Join us for a seminar with 박운영 on resume review and interview strategy. Event: Seminar · Date: Oct 15, 2026 · Time: 6:30 PM PST · Location: Seattle University · Speaker: 박운영 · Topics: Resume review, Interview strategy Map: https://www.google.com/maps/search/?api=1&query=Seattle+University

After: Join us for an inspiring seminar with 박운영 on career success! Seminar details: ▶ When: Oct 15, 2026 · 6:30 PM PST ▶ Where: Seattle University ▶ Speaker: 박운영 ▶ Price: Free What we'll cover: - Resume review - Interview strategy - Networking Register now and be the first to get expert tips! Register: https://pknic.org/events/career-1015

Before: every fact on one line, and no call to action

After: one fact a line, topics grouped, an ask and a link

X 280 characters, read in one glance

Before: PKNIC @pknic · 1h Join us for an engaging seminar with HR expert 박운영, where we'll dive into resume reviews and interview strategies. In this interactive session, HR professional 박운영 will guide you through resume review techniques and craft the perfect interview strategy. Don't miss out on this val

After: PKNIC @pknic · 1h Join us for an inspiring seminar with 박운영 on career advancement. 🗓 Oct 15, 2026 · 6:30 PM PST · Free 📍 Seattle University Register now and unlock your career potential! https://pknic.org/events/career-1015 #career #seminar #networking

Before: no date or place, and cut off mid-word

After: when and where first, built to fit

Move across each post to compare. The same brief and the same model, before and after, in real outputs.

05Results

From 30 minutes to under 5.

Posting one event to LinkedIn, Instagram and KakaoTalk took me at least 30 minutes: write, review, then check each post on the platform. Now I write the event details, pick the channels, check the previews and copy each post into its app, in under 5 minutes.

In the final Korean run across six channels, no post carried an invented date, price or venue.

Fast enough on 8GB

Early batches timed out. Timing each call showed that all of the time went to the model writing, and X was stuck in a loop, repeating one invented sentence until it ran out of room. X stopped asking for a body paragraph it never needed, every call got a token ceiling, and channels now run one at a time, fastest first.

Seconds per postBeforeAfter
KakaoTalk12.03.3
WhatsApp8.93.7
Xtimed out at 3004.1
Measured on the same machine.

06Looking back

What I would claim, and what comes next.

  1. 1

    One sample per cell

    Each model, channel and language ran once, at temperature 0. The patterns repeated across models, but the rankings are suggestive, not settled.

  2. 2

    The rubric is blind to facts

    Scores measure length, format and tone. I judged grounding by reading every output, which does not scale.

  3. 3

    Twelve human reviews so far

    The edit-and-reason loop that teaches the system its voice works end to end, with little data behind it yet.

  4. 4

    The scrub is strict on purpose

    It can delete a harmless sentence with a number, like “3 tips”. Every deletion is flagged so a reviewer sees what went.

Next

Image generation for event posters, and OAuth sign-in to each platform, so a post goes out on its own at a scheduled time.

All examples are unedited model output from the project's database and test runs, marked up by me. Dates and weekdays were checked against the calendar.

More work

Identity at national scale, AI in research budgeting, and the builds in between.

All work