All skills
ljagiello avatar

/ctf-misc

@61c2efe
by Lukasz Jagielloljagiello/ctf-skills3.4k stars
393

Provides miscellaneous CTF challenge techniques for problems that do not cleanly fit the main categories. Use for encoding puzzles, pyjails, bash jails, RF/SDR, DNS oddities, unicode tricks, esoteric languages, QR or audio puzzles, constraint solving, game theory, unusual sandbox escapes, and hybrid logic puzzles. Prefer a more specific skill first when the challenge is mainly web, pwn, reverse, forensics, malware, OSINT, or crypto. Treat this as the fallback skill for genuine cross-category or edge-case challenges, not the default starting point.

Use this Skill: https://skilld.dev/gh/ljagiello/ctf-skills/ctf-misc

This session only. Nothing lands on disk.

encodings.md

≈4.8k tokens on demand. Your agent reads this file only when SKILL.md points to it.

CTF Misc - Encodings & Media

Table of Contents

See also: encodings-advanced.md - Verilog/HDL, Gray code, binary tree encoding, RTF custom tags, SMS PDU decoding, multi-encoding solvers, UTF-9, pixel binary encoding, hex Sudoku + QR, TOPKEK, MaxiCode


Common Encodings

Base64

echo "encoded" | base64 -d
# Charset: A-Za-z0-9+/=

Base32

echo "OBUWG32DKRDHWMLUL53TI43OG5PWQNDSMRPXK3TSGR3DG3BRNY4V65DIGNPW2MDCGFWDGX3DGBSDG7I=" | base32 -d
# Charset: A-Z2-7= (no lowercase, no 0,1,8,9)

Hex

echo "68656c6c6f" | xxd -r -p

IEEE 754 Floating Point Encoding

Numbers that encode ASCII text when viewed as raw IEEE 754 bytes:

import struct

values = [240600592, 212.2753143310547, 2.7884192016691608e+23]

# Each float32 packs to 4 ASCII bytes
for v in values:
    packed = struct.pack('>f', v)  # Big-endian single precision
    print(f"{v} -> {packed}")      # b'Meta', b'CTF{', b'fl04'

# For double precision (8 bytes per value):
# struct.pack('>d', v)

Key insight: If challenge gives a list of numbers (mix of integers, decimals, scientific notation), try packing each as IEEE 754 float32 (struct.pack('>f', v)) — the 4 bytes often spell ASCII text.

UTF-16 Endianness Reversal (LACTF 2026)

Pattern (endians): Text "turned to Japanese" -- mojibake from UTF-16 endianness mismatch.

Fix: Reverse the encoding/decoding order:

# If encoded as UTF-16-LE but decoded as UTF-16-BE:
fixed = mojibake.encode('utf-16-be').decode('utf-16-le')

# If encoded as UTF-16-BE but decoded as UTF-16-LE:
fixed = mojibake.encode('utf-16-le').decode('utf-16-be')

Identification: Text appears as CJK characters (Japanese/Chinese), challenge mentions "translation" or "endian".

BCD (Binary-Coded Decimal) Encoding (VuwCTF 2025)

Pattern: Challenge name hints at ratio (e.g., "1.5x" = 1.5:1 byte ratio). Each nibble encodes one decimal digit.

def bcd_decode(data):
    """Decode BCD: each byte = 2 decimal digits."""
    return ''.join(f'{(b>>4)&0xf}{b&0xf}' for b in data)

# Then convert decimal string to ASCII
ascii_text = ''.join(chr(int(decoded[i:i+2])) for i in range(0, len(decoded), 2))

Multi-Layer Encoding Detection (0xFun 2026)

Pattern (139 steps): Recursive decoding with troll flags as decoys.

Critical rule: When data is all hex chars (0-9, a-f), decode as hex FIRST, not base64 (which also accepts those chars). Distinguish compression by magic: zlib 78 01/78 9c, gzip 1f 8b, bz2 42 5a 68 (BZh), lzma/xz fd 37 7a 58 5a 00.

import base64, zlib, gzip, bz2, lzma

def try_xor(data, key_hint='flag'):
    """Single-byte brute 256 + repeating-key hint; xor_flag heuristic."""
    raw = data.encode() if isinstance(data, str) else data
    best = None
    # First pass: exact hint match (case-sensitive) — avoids FLAG false positive
    # (e.g. key 0x20 maps FLAG->flag; exact pass prefers the true key).
    for key in range(256):
        decoded = bytes(b ^ key for b in raw)
        if key_hint and decoded.count(key_hint.encode()) > 0:
            return decoded
    # Second pass: case-insensitive flag + printable fallback
    for key in range(256):
        decoded = bytes(b ^ key for b in raw)
        printable = all(32 <= b < 127 or b in b'\n\r\t' for b in decoded[:64]) if decoded else False
        if printable and b'flag' in decoded.lower():
            return decoded
        if printable and best is None:
            best = decoded
    # repeating-key hint (key_hint as repeating key)
    if key_hint:
        kh = key_hint.encode()
        decoded = bytes(b ^ kh[i % len(kh)] for i, b in enumerate(raw))
        if b'flag{' in decoded or b'CTF{' in decoded:
            return decoded
    return best

def _as_text(raw: bytes):
    """Decode only for str-only passes; return None if not text-safe."""
    try:
        return raw.decode('ascii')
    except UnicodeDecodeError:
        return None

def auto_decode(data):
    # Keep raw bytes across passes — never lossy-decode mid-chain with
    # errors='replace' (that corrupts the next pass's input). Text-only
    # decoders (hex/base64) run on the ASCII view; compression/xor run on bytes.
    raw = data.encode() if isinstance(data, str) else bytes(data)
    while True:
        text = _as_text(raw)
        s = text.strip() if text is not None else None
        if s is not None and s.startswith('REAL_DATA_FOLLOWS:'):
            raw = s.split(':', 1)[1].encode()
            continue
        # Distinguish zlib (78 01/78 9c) vs gzip (1f 8b) vs bz2 vs lzma magic before text decoders
        if raw.startswith(b'\x78\x01') or raw.startswith(b'\x78\x9c'):  # zlib
            try:
                raw = zlib.decompress(raw)
                continue
            except Exception:
                pass
        if raw.startswith(b'\x1f\x8b'):  # gzip 1f 8b
            try:
                raw = gzip.decompress(raw)
                continue
            except Exception:
                pass
        if raw.startswith(b'BZh'):  # bz2
            try:
                raw = bz2.decompress(raw)
                continue
            except Exception:
                pass
        if raw.startswith(b'\xfd7zXZ\x00'):  # lzma/xz
            try:
                raw = lzma.decompress(raw)
                continue
            except Exception:
                pass
        # xor branch — try_xor(data, key_hint='flag') single-byte brute 256 + repeating-key hint
        # (two-pass: exact hint first, then case-insensitive printable fallback)
        xor_res = try_xor(raw, key_hint='flag')
        if xor_res is not None and xor_res != raw and b'flag' in xor_res.lower():
            raw = xor_res
            continue
        # Prioritize hex when ambiguous — hex-first rule (str-only passes)
        if s is not None and len(s) % 2 == 0 and len(s) > 0 and all(c in '0123456789abcdefABCDEF' for c in s):
            try:
                raw = bytes.fromhex(s)
                continue
            except Exception:
                pass
        if s is not None and len(s) % 4 == 0 and set(s) <= set('ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/='):
            try:
                raw = base64.b64decode(s)
                continue
            except Exception:
                pass
        break
    return raw

Ignore troll flags — check for "keep decoding" or "REAL_DATA_FOLLOWS:" markers.

CyberChef Magic verify: paste remaining bytes into CyberChef Magic operation — it brute-forces the same xor/base/compression stack and confirms the manual auto_decode chain. Use Magic's intensity slider to replicate the hex-first rule.

URL Encoding

import urllib.parse
urllib.parse.unquote('hello%20world')

ROT13 / Caesar

echo "uryyb" | tr 'a-zA-Z' 'n-za-mN-ZA-M'

ROT13 patterns: gur = "the", synt = "flag"

Caesar Brute Force

text = "Khoor Zruog"
for shift in range(26):
    decoded = ''.join(
        chr((ord(c) - 65 - shift) % 26 + 65) if c.isupper()
        else chr((ord(c) - 97 - shift) % 26 + 97) if c.islower()
        else c for c in text)
    print(f"{shift:2d}: {decoded}")

QR Codes

Basic Commands

zbarimg qrcode.png           # Decode
zbarimg -S*.enable qr.png    # All barcode types
qrencode -o out.png "data"   # Encode

QR Structure

Finder patterns (3 corners): 7x7 modules at top-left, top-right, bottom-left

Version formula: (version * 4) + 17 modules per side

Repairing Damaged QR

from PIL import Image
import numpy as np

img = Image.open('damaged_qr.png')
arr = np.array(img)

# Convert to binary
gray = np.mean(arr, axis=2)
binary = (gray < 128).astype(int)

# Find QR bounds
rows = np.any(binary, axis=1)
cols = np.any(binary, axis=0)
rmin, rmax = np.where(rows)[0][[0, -1]]
cmin, cmax = np.where(cols)[0][[0, -1]]

# Check finder patterns
qr = binary[rmin:rmax+1, cmin:cmax+1]
print("Top-left:", qr[0:7, 0:7].sum())  # Should be ~25

Finder Pattern Template

finder_pattern = [
    [1,1,1,1,1,1,1],
    [1,0,0,0,0,0,1],
    [1,0,1,1,1,0,1],
    [1,0,1,1,1,0,1],
    [1,0,1,1,1,0,1],
    [1,0,0,0,0,0,1],
    [1,1,1,1,1,1,1],
]

QR Code Chunk Reassembly (LACTF 2026)

Pattern (error-correction): QR code split into grid of chunks (e.g., 5x5 of 9x9 pixels), shuffled.

Solving approach:

  1. Fix known chunks: Use structural patterns -- finder patterns (3 corners), timing patterns, alignment patterns -- to place ~50% of chunks
  2. Extract codeword constraints: For each candidate payload length, use QR spec to identify which pixels are invariant across encodings
  3. Backtracking search: Assign remaining chunks under pixel constraints until QR decodes successfully

Tools: segno (Python QR library), zbarimg for decoding.

QR Code Chunk Reassembly via Indexed Directories (UTCTF 2026)

Pattern (QRecreate): QR code split into numbered chunks stored in separate directories. Directory names encode the chunk index as base64 (e.g., MDAx → 001 → index 1).

Solving approach:

  1. Decode each directory name from base64 to get the numeric index
  2. Sort chunks by decoded index
  3. Arrange in a grid (e.g., 100 chunks → 10x10) and stitch into a single image
  4. Decode the reconstructed QR code
import os, base64, math
from PIL import Image

# 1. Decode directory names to get indices
chunks = []
for dirname in os.listdir('chunks/'):
    index = int(base64.b64decode(dirname).decode())
    tile = Image.open(f'chunks/{dirname}/tile.png')
    chunks.append((index, tile))

# 2. Sort by index and arrange in grid
chunks.sort(key=lambda x: x[0])
n = len(chunks)
side = int(math.isqrt(n))
tile_w, tile_h = chunks[0][1].size

canvas = Image.new("RGB", (side * tile_w, side * tile_h), (255, 255, 255))
for i, (_, tile) in enumerate(chunks):
    r, c = divmod(i, side)
    canvas.paste(tile, (c * tile_w, r * tile_h))

canvas.save('reconstructed_qr.png')
# 3. Decode with zbarimg or pyzbar

Key insight: Unlike the LACTF variant (shuffled chunks requiring structural analysis), indexed chunks just need sorting. The challenge is recognizing that directory names are base64-encoded indices. Check base64 -d on folder names when they look like random strings.


Multi-Stage URL Encoding Chain (UTCTF 2026)

Pattern (Breadcrumbs): Flag is hidden behind a chain of URLs, each encoded differently. Follow the breadcrumbs across external resources (GitHub Gists, Pastebin, etc.), decoding at each hop.

Common encoding layers per hop:

  1. Base64 → URL to next resource
  2. Hex → URL to next resource (e.g., 68747470733a2f2f... = https://...)
  3. ROT13 → final flag

Decoding workflow:

import base64, codecs

# Hop 1: Base64
hop1 = "aHR0cHM6Ly9naXN0Lmdp..."
url2 = base64.b64decode(hop1).decode()

# Hop 2: Hex-encoded URL
hop2 = "68747470733a2f2f..."
url3 = bytes.fromhex(hop2).decode()

# Hop 3: ROT13-encoded flag
hop3 = "hgsynt{...}"
flag = codecs.decode(hop3, 'rot_13')

Key insight: Each resource contains a hint about the next encoding (e.g., "Three letters follow" hints at 3-character encoding like hex). Look for contextual clues in surrounding text (poetry, comments, filenames) that indicate the encoding type.

Detection: Challenge mentions "trail", "breadcrumbs", "follow", or "scavenger hunt". First resource contains what looks like encoded data rather than a direct flag.


Esoteric Languages

Language Pattern
Brainfuck ++++++++++[>+++++++>
Whitespace Only spaces, tabs, newlines (or S/T/L substitution)
Ook! Ook. Ook? Ook!
Malbolge Extremely obfuscated
Piet Image-based

Whitespace Language Parser (BYPASS CTF 2025)

Pattern (Whispers of the Cursed Scroll): File contains only S (space), T (tab), L (linefeed) characters — or visible substitutes. Stack-based virtual machine (VM) with PUSH, OUTPUT, and EXIT instructions.

Instruction set (IMP = Instruction Modification Parameter):

Instruction Encoding Action
PUSH S S + sign + binary + L Push number to stack (S=0, T=1, L=terminator)
OUTPUT CHAR T L S S Pop stack, print as ASCII character
EXIT L L L Halt program
def solve_whitespace(content):
    # Convert to S/T/L tokens (handle both raw whitespace and visible chars)
    if any(c in content for c in 'STL'):
        code = [c for c in content if c in 'STL']
    else:
        code = [{'\\s': 'S', '\\t': 'T', '\\n': 'L'}.get(c, '') for c in content]
        code = [c for c in code if c]

    stack, output, i = [], "", 0

    while i < len(code):
        if code[i:i+2] == ['S', 'S']:  # PUSH
            i += 2
            sign = 1 if code[i] == 'S' else -1
            i += 1
            val = 0
            while i < len(code) and code[i] != 'L':
                val = (val << 1) + (1 if code[i] == 'T' else 0)
                i += 1
            i += 1  # skip terminator L
            stack.append(sign * val)
        elif code[i:i+4] == ['T', 'L', 'S', 'S']:  # OUTPUT CHAR
            i += 4
            if stack:
                output += chr(stack.pop())
        elif code[i:i+3] == ['L', 'L', 'L']:  # EXIT
            break
        else:
            i += 1

    return output

Identification: File with only whitespace characters, or challenge mentions "invisible code", "blank page", or uses S/T/L substitution. Try Whitespace interpreter online for quick testing.


Custom Brainfuck Variants (Themed Esolangs)

Pattern: File contains repetitive themed words (e.g., "arch", "linux", "btw") used as substitutes for Brainfuck operations. Common in Easy/Misc CTF challenges.

Identification:

  • File is ASCII text with very long lines of repeated words
  • Small vocabulary (5-8 unique words)
  • One word appears as a line terminator (maps to . output)
  • Two words are used for increment/decrement (one has many repeats per line)
  • Words often relate to a meme or theme (e.g., "I use Arch Linux BTW")

Standard Brainfuck operations to map:

Op Meaning Typical pattern
+ Increment cell Most repeated word (defines values)
- Decrement cell Second most repeated word
> Move pointer right Short word, appears alone or with .
< Move pointer left Paired with > word
[ Begin loop Appears at start of lines with ] counterpart
] End loop Appears at end of same lines as [
. Output char Line terminator word

Solving approach:

from collections import Counter
words = content.split()
freq = Counter(words)
# Most frequent = likely + or -, line-ender = likely .

# Map words to BF ops, translate, run standard BF interpreter
mapping = {'arch': '+', 'linux': '-', 'i': '>', 'use': '<',
           'the': '[', 'way': ']', 'btw': '.'}
bf = ''.join(mapping.get(w, '') for w in words)
# Then execute bf string with a standard Brainfuck interpreter

Real example (0xL4ugh CTF - "iUseArchBTW"): .archbtw extension, "I use Arch Linux BTW" meme theme.

Tips: Try swapping +/- or >/< if output is not ASCII. Verify output starts with known flag format.


Multi-Layer Esoteric Language Chains (Break In 2016)

Challenges may stack multiple esoteric languages requiring sequential interpretation:

  1. Piet: Visual programming language using colored pixel blocks. Execute PNG images as code:
npiet challenge.png         # npiet interpreter
# Or: java -jar PietDev.jar challenge.png
  1. Malbolge: Extremely difficult esoteric language. Decode output from previous layer:
# Piet output → base64 decode → Malbolge source
echo "piet_output" | base64 -d > program.mal
malbolge program.mal        # Or use online interpreter

Common esoteric chains: Piet → base64 → Malbolge, Brainfuck → Ook → Whitespace, JSFuck → standard JS.

Key insight: When a PNG file doesn't contain obvious visual stego, try interpreting it as Piet code. Use file + visual inspection to identify the first layer, then decode sequentially.


base65536 CJK Unicode Binary Encoding (IceCTF 2018)

Pattern: A blob that looks like a wall of Chinese characters (CJK Unified Ideographs) is actually a base65536 encoding: each character carries two bytes of data, mapping 0x0000..0xFFFF to a picked subset of 65,536 Unicode codepoints. Detect by file reporting "Unicode text, UTF-8" with mostly CJK codepoints; decode with the base65536 npm package or the Python port.

# Node.js / npm path
npm install -g base65536
echo -n "宝䀈䀋..." | base65536 --decode > out.bin

# Python port
pip install base65536
python3 - <<'PY'
import base65536, sys
sys.stdout.buffer.write(base65536.decode(open("blob.txt").read()))
PY > out.bin

file out.bin
# common outcome: "Zip archive data" or "ELF 64-bit"

Key insight: base64 expands 3 bytes → 4 chars; base65536 expands 2 bytes → 1 Unicode codepoint, and since a codepoint renders as 1–4 UTF-8 bytes the encoded stream actually expands by ~2× on disk — but visually it looks compact, which is the CTF trick. Any wall of Unicode that lacks variance across the Basic Multilingual Plane and is dominated by CJK, Hangul, or Tibetan is a candidate. Also check base1024 (BMP), base2048, base4096, and base32768 for related tricks.

References: IceCTF 2018 — Rabbit Hole, writeup 11421

Source: SKILL.md on GitHub

3 alerts17d5 checks · Risk CRITICAL
  • Gen Agent Trust Hub17d

    The skill is a comprehensive reference and cheat sheet repository for Capture The Flag (CTF) miscellaneous challenges, including jail escapes, cryptography puzzles, and Linux privilege escalation techniques. While automated scanners flag signatures of exploit payloads and shell patterns within the documentation, these are entirely educational examples and static reference material aligned with the primary purpose of the skill.

  • Socket17d

    2 alerts: gptSecurity, gptMalware

  • Snyk17d

    Risk: HIGH · 1 issue

  • Runlayer6mo

    5/7 files flagged

  • ZeroLeaks5mo

    2 findings · Score: 80/100

Signed by skilld at 61c2efe. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 weeks ago.

Activeupdated 3 weeks ago
metadata
{
  "user-invocable": "false"
}
All 1 allowed tools
Bash Read Write Edit Glob Grep Task WebFetch WebSearch Skill
Other metadata
compatibility
Requires filesystem-based agent (Claude Code or similar) with bash, Python 3, and internet access for tool installation.

README badge

README badge for ljagiello/ctf-skills/ctf-misc