All skills
inference-shell avatar

/python-sdk

@fbe0aa4

Python SDK for inference.sh - run AI apps, build agents, and integrate with all models. Package: inferencesh (pip install inferencesh). Supports sync/async, streaming, file uploads. Build agents with template or ad-hoc patterns, tool builder API, skills, and human approval. Use for: Python integration, AI apps, agent development, RAG pipelines, automation. Triggers: python sdk, inferencesh, pip install, python api, python client, async inference, python agent, tool builder python, programmatic ai, python integration, sdk python

Use this Skill: https://skilld.dev/gh/inference-shell/skills/python-sdk

This session only. Nothing lands on disk.

referencesstreaming.md

≈1.5k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Streaming Reference

Real-time progress updates and Server-Sent Events (SSE) handling.

Task Status Flow

RECEIVED (1) → QUEUED (2) → SCHEDULED (3) → PREPARING (4)
→ SERVING (5) → SETTING_UP (6) → RUNNING (7) → UPLOADING (8)
→ COMPLETED (10), FAILED (11), or CANCELLED (12)

Basic Streaming

from inferencesh import inference

client = inference(api_key="inf_...")

for update in client.run({
    "app": "google/veo-3-1-fast",
    "input": {"prompt": "A sunset timelapse"}
}, stream=True):
    print(f"Status: {update['status']}")

Handling Different Update Types

for update in client.run(config, stream=True):
    status = update.get("status")

    # Task state changes
    if status == "queued":
        print("Task queued, waiting for worker...")
    elif status == "running":
        print("Task is running...")
    elif status == "completed":
        print("Done!")
        print(f"Output: {update.get('output')}")
    elif status == "failed":
        print(f"Error: {update.get('error')}")

    # Progress logs
    if update.get("logs"):
        for log in update["logs"]:
            print(f"  Log: {log}")

    # Partial outputs
    if update.get("partial_output"):
        print(f"  Partial: {update['partial_output']}")

Progress Tracking with UI

import sys

def progress_bar(current, total, width=50):
    filled = int(width * current / total)
    bar = "█" * filled + "░" * (width - filled)
    percent = current / total * 100
    sys.stdout.write(f"\r[{bar}] {percent:.1f}%")
    sys.stdout.flush()

for update in client.run(config, stream=True):
    if update.get("progress"):
        progress_bar(update["progress"]["current"], update["progress"]["total"])

    if update.get("status") == "completed":
        print("\n✓ Complete!")

Streaming with Timeout

import time

start = time.time()
timeout = 300  # 5 minutes

for update in client.run(config, stream=True):
    if time.time() - start > timeout:
        print("Timeout reached")
        break

    print(f"Status: {update['status']}")

    if update.get("status") in ["completed", "failed"]:
        break

Async Streaming

from inferencesh import async_inference
import asyncio

async def stream_task():
    client = async_inference(api_key="inf_...")

    async for update in client.run({
        "app": "google/veo-3-1-fast",
        "input": {"prompt": "Ocean waves"}
    }, stream=True):
        print(f"Status: {update['status']}")

        if update.get("status") == "completed":
            return update.get("output")

result = asyncio.run(stream_task())

Agent Streaming

agent = client.agent("my-org/assistant@latest")

def on_message(msg):
    if msg.get("content"):
        # Stream text as it arrives
        print(msg["content"], end="", flush=True)

    if msg.get("type") == "thinking":
        print(f"\n[Thinking: {msg.get('content')}]")

def on_tool_call(call):
    print(f"\n[Calling tool: {call.name}]")
    result = execute_tool(call.name, call.args)
    agent.submit_tool_result(call.id, result)

response = agent.send_message(
    "Explain quantum entanglement",
    on_message=on_message,
    on_tool_call=on_tool_call
)

Reconnection Handling

from inferencesh import inference, StreamingOptions

client = inference(api_key="inf_...")

options = StreamingOptions(
    max_retries=3,
    retry_delay=1.0,  # seconds
    chunk_size=1024
)

for update in client.run(config, stream=True, options=options):
    print(update)

Multiple Streams in Parallel

from inferencesh import async_inference
import asyncio

async def run_parallel():
    client = async_inference(api_key="inf_...")

    configs = [
        {"app": "infsh/flux-1-dev", "input": {"prompt": "A mountain"}},
        {"app": "infsh/flux-1-dev", "input": {"prompt": "An ocean"}},
        {"app": "infsh/flux-1-dev", "input": {"prompt": "A forest"}}
    ]

    async def stream_one(config, index):
        async for update in client.run(config, stream=True):
            print(f"[{index}] {update['status']}")
            if update.get("status") == "completed":
                return update.get("output")

    results = await asyncio.gather(*[
        stream_one(c, i) for i, c in enumerate(configs)
    ])
    return results

results = asyncio.run(run_parallel())

Cancelling a Stream

task_id = None

try:
    for update in client.run(config, stream=True):
        task_id = update.get("id")
        print(f"Status: {update['status']}")

        if should_cancel():
            break
finally:
    if task_id:
        client.cancel_task(task_id)
        print("Task cancelled")

Collecting All Logs

all_logs = []

for update in client.run(config, stream=True):
    if update.get("logs"):
        all_logs.extend(update["logs"])

    if update.get("status") == "completed":
        print("Final logs:")
        for log in all_logs:
            print(f"  {log}")

Custom Stream Processing

class StreamProcessor:
    def __init__(self):
        self.logs = []
        self.start_time = None
        self.end_time = None

    def process(self, update):
        if self.start_time is None:
            self.start_time = time.time()

        if update.get("logs"):
            self.logs.extend(update["logs"])

        if update.get("status") in ["completed", "failed"]:
            self.end_time = time.time()
            return True  # Done

        return False  # Continue

    @property
    def duration(self):
        if self.start_time and self.end_time:
            return self.end_time - self.start_time
        return None

processor = StreamProcessor()

for update in client.run(config, stream=True):
    if processor.process(update):
        break

print(f"Duration: {processor.duration:.2f}s")
print(f"Logs: {len(processor.logs)}")

Source: SKILL.md on GitHub

3 warnings16d5 checks · Risk MEDIUM
  • Gen Agent Trust Hub16d

    The skill provides a Python SDK for interacting with the inference.sh platform. While intended for development, it contains documentation examples that promote insecure coding practices, such as using `eval()` on unsanitized tool arguments. It also features automatic file upload capabilities that could be exploited to exfiltrate sensitive local files if an agent is tricked into processing malicious paths.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer6mo

    1/7 files flagged

  • ZeroLeaks5mo

    3 findings · Score: 69/100

Signed by skilld at fbe0aa4. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 3 months ago
What it can do
Runs commands
All 2 allowed tools
Bash(pip install inferencesh)Bash(python *)
  • Python
  • sdk
  • inference-sh
  • ai-agents
  • llm
  • async
  • streaming
  • tool-builder
  • rag
  • api-client

README badge

README badge for inference-shell/skills/python-sdk

Provides a Python SDK for inference.sh that runs AI apps, builds agents with tool builders and human approval workflows, and integrates 250+ models via a single API. Supports sync/async execution, streaming, file uploads, and stateful sessions—targets Python developers building AI applications and agent systems.

Generated from the current SKILL.md.

What Python versions does this SDK support?
Python 3.8 and later. Install with `pip install inferencesh` for sync support or `pip install inferencesh[async]` for async/await.
Does this support async/await?
Yes. Use `async_inference` instead of `inference` and await calls like `await client.run(...)` and `await agent.send_message(...)`.
Can I build agents with custom tools?
Yes. Use the tool builder API to define client tools, app tools (calls to inference.sh apps), agent tools (delegation to sub-agents), and webhook tools (external APIs). Tools can require human approval before execution.
What models are available as core apps for agents?
Claude Sonnet 4, Claude 3.5 Haiku, GPT-4o, and GPT-4o Mini. Access via app references like `infsh/claude-sonnet-4@latest`.
Does the SDK handle file uploads?
Yes. Files can be auto-uploaded by passing file paths directly in input, or manually uploaded via `client.upload_file()` with optional metadata like custom filename and content-type.

Generated from the current SKILL.md. These answers refresh after source changes.