Model appraisal — August 22, 2026
DeepSWE Value on a 1M-Token Codebase: Flash, Pro, or Luna?
A costed comparison of DeepSeek V4 Flash 0731, V4 Pro, and GPT-5.6 Luna that counts the extra effort tokens a large-codebase agent actually spends.
Read the appraisal →
Research Notes — July 30, 2026
The State of Bank Refresh Tokens in Canada — and What Changes When the Consumer-Driven Banking Act Lands
An experiment scanning Plaid's 198 Canadian institutions found only 2 use OAuth. Every Big-5 bank is credential-scraping. Canada's new open-banking law will change this — here's the timeline and what actually improves.
Read the research →
Article Review — July 14, 2026
Five Signals a Workflow Is Broken — and Ripe for AI
Tabs, copy/paste, waiting, rework, handoffs — a field-ready test for spotting the workflows worth automating, before you build a thing.
Read the review →
Article Review — July 14, 2026
Kalshi Bets on a Market for AI Computing Power
Kalshi launched a forward curve for compute. What a public price on AI hardware means for anyone paying an LLM token bill.
Read the review →
Model appraisals post here too — each new LLM gets its tokenizer, thinking-token, and speed measured live against our stack. Methodology and datasets live in the Benchmarks section and the LLM-Cost-Comparison repo.
