Articles
6 long-form essays. Deep dives, opinions, tutorials.
Ref PG-ART-006 Date 2026-07-03 Category LLM Read 7 min
Claude Code hid a watermark that flagged Chinese users. Chinese models write weaker code for Americans. If you cannot diff your stack, it can decide who you are.
Read article > Ref PG-ART-005 Date 2026-05-01 Category Context Read 7 min
Bigger is not better. The advertised number lies. Pick the smallest window that fits the work and you save money, time, and accuracy.
Read article > Ref PG-ART-004 Date 2026-04-30 Category Llama Cpp Read 4 min
43 words a second. 0.29 second wait. 20 of 20 coding answers correct. 19.5 GB used out of 24. The card runs this AI very well — and it can hold a 128k window with 4 chats at the same time.
Read article > Ref PG-ART-003 Date 2026-04-29 Category Llama Cpp Read 5 min
44 words a second. 0.26 second wait. 19 of 20 coding answers worked. 18 GB used out of 24. The card runs this AI well — and it can hold a 128k window with 4 chats at the same time.
Read article > Ref PG-ART-002 Date 2026-04-22 Category VLLM Read 5 min
105 words a second solo. 160 words a second with 2 chats. 0.56 second wait. 32 GB card almost full. Tested on vLLM.
Read article > Ref PG-ART-001 Date 2026-04-16 Category Claude Read 3 min
Three ways to work with Claude. They are not interchangeable. Using the wrong one wastes hours. Here is when to reach for each.
Read article >