All skills
oaustegard avatar

/declauding

@6777f98

Load at the start of any task whose deliverable is prose another person will read — a PR description, commit message, README, doc, postmortem, blog post, issue or review comment, report, release note or essay — before drafting it, whether the result is pushed, posted, published or saved to a file, without being asked. Draft, then run this pass on the draft before handing it over. Also use when someone says "de-claude", "de-slop", "humanize this", "this reads like AI", "make it sound human", or asks for a voice, tone or register edit. Rewrites the constructions that mark prose as model-written (staged reveals, verdict headers, aphoristic closers, "it's not X, it's Y", em-dash drama, forced triads, flat-certainty adverbs) into plain technical prose and checks the rewrite kept every claim. Not for fiction, poetry, code, or quoted text; for a full adversarial review of a deliverable use challenging.

  • 23 files
  • 249 KB
  • Updated last week
  • GitHub

Use this Skill: https://skilld.dev/gh/oaustegard/claude-skills/declauding

This session only. Nothing lands on disk.

testssample-clean.md

≈431 tokens on demand. Your agent reads this file only when SKILL.md points to it.

A few words on DS4

I didn't expect DwarfStar 4 to become so popular so fast. It is clear that there was a need for single-model integration focused local AI experience, and that a few things happened together: the release of a quasi-frontier model that is large and fast enough to change the game of local inference, and the fact that it works extremely well with an extremely asymmetric quants recipe of 2/8 bit, so that 96 or 128GB of RAM are enough to run it.

The last week was funny and also tiring, I worked 14 hours per day on average. My normal average is 4/6 since early Redis times, but the first few months of Redis were like that.

So, what's next? Is this a project that starts and ends with DeepSeek v4 Flash? Nope, the model can change over time. The space will be occupied, in my vision, by the best current open weights model that is practically fast on a high end Mac or GPU in a box gear.

Power capping

Long local inference runs can keep the GPU busy for extended periods. If you care more about heat, fan noise, battery life on MacBooks, or reducing thermal stress on the hardware than about maximum throughput, use --power N. --power 100 is the default and means full speed. Lower values ask DwarfStar to target that percentage of GPU usage. DwarfStar does this by measuring GPU work time and inserting small sleeps between work units: during prefill it sleeps between layers, and during generation it sleeps between decoded tokens. This reduces sustained load without changing model output.

I did not expect the overlap to survive a 3x range in bits per weight and two different quant families.

"You took eleven seconds," Sonnet said.

"I know."

"On four sentences."

"Then take twelve."

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at 6777f98. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last week
metadata
{
  "version": "0.9.3"
}

README badge

README badge for oaustegard/claude-skills/declauding