All skills
inference-shell avatar

/python-sdk

@fbe0aa4

Python SDK for inference.sh - run AI apps, build agents, and integrate with all models. Package: inferencesh (pip install inferencesh). Supports sync/async, streaming, file uploads. Build agents with template or ad-hoc patterns, tool builder API, skills, and human approval. Use for: Python integration, AI apps, agent development, RAG pipelines, automation. Triggers: python sdk, inferencesh, pip install, python api, python client, async inference, python agent, tool builder python, programmatic ai, python integration, sdk python

Use this Skill: https://skilld.dev/gh/inference-shell/skills/python-sdk

This session only. Nothing lands on disk.

referencesagent-patterns.md

≈1.8k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Agent Patterns

Common patterns for building agents with the Python SDK.

Multi-Agent Orchestration

Delegate tasks to specialized sub-agents:

from inferencesh import inference, agent_tool, string

client = inference(api_key="inf_...")

# Define sub-agents as tools
researcher = (
    agent_tool("research", "my-org/researcher@latest")
    .describe("Research a topic thoroughly")
    .param("topic", string("Topic to research"))
    .build()
)

writer = (
    agent_tool("write", "my-org/writer@latest")
    .describe("Write content based on research")
    .param("outline", string("Content outline"))
    .param("research", string("Research findings"))
    .build()
)

# Orchestrator agent
orchestrator = client.agent({
    "core_app": {"ref": "infsh/claude-sonnet-4@latest"},
    "system_prompt": """You are an orchestrator that:
1. Uses the research tool to gather information
2. Uses the write tool to create content
Coordinate between agents to produce high-quality output.""",
    "tools": [researcher, writer]
})

response = orchestrator.send_message("Create a blog post about AI agents")

RAG Pattern (Retrieval-Augmented Generation)

Combine search with LLM responses:

from inferencesh import inference, app_tool, string

client = inference(api_key="inf_...")

# Search tool
search = (
    app_tool("search", "tavily/search-assistant@latest")
    .describe("Search the web for current information")
    .param("query", string("Search query"))
    .build()
)

# RAG agent
rag_agent = client.agent({
    "core_app": {"ref": "infsh/claude-sonnet-4@latest"},
    "system_prompt": """You help users with current information.
When asked about recent events or facts you're unsure about,
use the search tool to find accurate, up-to-date information.
Always cite your sources.""",
    "tools": [search]
})

response = rag_agent.send_message("What are the latest developments in quantum computing?")

Code Execution Pattern

Agents that can write and run code:

from inferencesh import inference, internal_tools

client = inference(api_key="inf_...")

config = (
    internal_tools()
    .code_execution(True)
    .build()
)

coder = client.agent({
    "core_app": {"ref": "infsh/claude-sonnet-4@latest"},
    "system_prompt": """You are a Python coding assistant.
Write code to solve problems and execute it to verify it works.
Explain your approach and show the output.""",
    "internal_tools": config
})

response = coder.send_message("Calculate the first 20 Fibonacci numbers")

Human-in-the-Loop Pattern

Require approval for sensitive operations:

from inferencesh import inference, tool, string

client = inference(api_key="inf_...")

# Tool requiring approval
delete_file = (
    tool("delete_file")
    .describe("Delete a file from the filesystem")
    .param("path", string("File path to delete"))
    .require_approval()
    .build()
)

def handle_tool(call):
    if call.requires_approval:
        print(f"\n⚠️  Agent wants to: {call.name}")
        print(f"   Arguments: {call.args}")
        confirm = input("Allow? (y/n): ")

        if confirm.lower() == 'y':
            result = execute_operation(call.name, call.args)
            agent.submit_tool_result(call.id, result)
        else:
            agent.submit_tool_result(call.id, {
                "error": "Operation denied by user"
            })
    else:
        result = execute_operation(call.name, call.args)
        agent.submit_tool_result(call.id, result)

agent = client.agent({
    "core_app": {"ref": "infsh/claude-sonnet-4@latest"},
    "tools": [delete_file]
})

response = agent.send_message(
    "Clean up temporary files in /tmp/myapp",
    on_tool_call=handle_tool
)

Conversation Memory Pattern

Maintain context across sessions:

import json
from inferencesh import inference

client = inference(api_key="inf_...")

def save_chat(agent, filepath):
    chat = agent.get_chat()
    with open(filepath, 'w') as f:
        json.dump(chat, f)

def load_chat(agent, filepath):
    try:
        with open(filepath, 'r') as f:
            chat = json.load(f)
            # Restore conversation by replaying messages
            for msg in chat['messages']:
                if msg['role'] == 'user':
                    agent.send_message(msg['content'])
    except FileNotFoundError:
        pass

agent = client.agent("my-org/assistant@latest")

# Load previous conversation
load_chat(agent, "conversation.json")

# Continue conversation
response = agent.send_message("Continue where we left off")

# Save for next session
save_chat(agent, "conversation.json")

Streaming with Progress UI

Real-time updates for better UX:

from inferencesh import inference
import sys

client = inference(api_key="inf_...")
agent = client.agent("my-org/assistant@latest")

def stream_handler(msg):
    if msg.get("content"):
        sys.stdout.write(msg["content"])
        sys.stdout.flush()

def tool_handler(call):
    print(f"\n🔧 Using tool: {call.name}")
    # Execute and return result
    result = execute_tool(call.name, call.args)
    agent.submit_tool_result(call.id, result)
    print("✅ Tool completed")

response = agent.send_message(
    "Generate a report on market trends",
    on_message=stream_handler,
    on_tool_call=tool_handler
)

print("\n\n📊 Report complete!")

Error Recovery Pattern

Graceful handling of failures:

from inferencesh import inference, RequirementsNotMetException
import time

client = inference(api_key="inf_...")

def robust_run(config, max_retries=3):
    for attempt in range(max_retries):
        try:
            return client.run(config)
        except RequirementsNotMetException as e:
            print(f"Missing requirements: {e.errors}")
            raise
        except RuntimeError as e:
            if attempt < max_retries - 1:
                wait = 2 ** attempt
                print(f"Error: {e}. Retrying in {wait}s...")
                time.sleep(wait)
            else:
                raise

result = robust_run({
    "app": "infsh/flux-1-dev",
    "input": {"prompt": "A serene landscape"}
})

Batch Processing Pattern

Process multiple items efficiently:

from inferencesh import async_inference
import asyncio

async def process_batch(items):
    client = async_inference(api_key="inf_...")

    async def process_one(item):
        result = await client.run({
            "app": "infsh/flux-1-dev",
            "input": {"prompt": item}
        })
        return result

    # Process in parallel with concurrency limit
    semaphore = asyncio.Semaphore(5)  # Max 5 concurrent

    async def bounded_process(item):
        async with semaphore:
            return await process_one(item)

    results = await asyncio.gather(*[
        bounded_process(item) for item in items
    ])
    return results

prompts = [
    "A mountain sunrise",
    "A city at night",
    "An ocean sunset",
    "A forest path"
]

results = asyncio.run(process_batch(prompts))

Source: SKILL.md on GitHub

3 warnings16d5 checks · Risk MEDIUM
  • Gen Agent Trust Hub16d

    The skill provides a Python SDK for interacting with the inference.sh platform. While intended for development, it contains documentation examples that promote insecure coding practices, such as using `eval()` on unsanitized tool arguments. It also features automatic file upload capabilities that could be exploited to exfiltrate sensitive local files if an agent is tricked into processing malicious paths.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer6mo

    1/7 files flagged

  • ZeroLeaks5mo

    3 findings · Score: 69/100

Signed by skilld at fbe0aa4. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 3 months ago
What it can do
Runs commands
All 2 allowed tools
Bash(pip install inferencesh)Bash(python *)
  • Python
  • sdk
  • inference-sh
  • ai-agents
  • llm
  • async
  • streaming
  • tool-builder
  • rag
  • api-client

README badge

README badge for inference-shell/skills/python-sdk

Provides a Python SDK for inference.sh that runs AI apps, builds agents with tool builders and human approval workflows, and integrates 250+ models via a single API. Supports sync/async execution, streaming, file uploads, and stateful sessions—targets Python developers building AI applications and agent systems.

Generated from the current SKILL.md.

What Python versions does this SDK support?
Python 3.8 and later. Install with `pip install inferencesh` for sync support or `pip install inferencesh[async]` for async/await.
Does this support async/await?
Yes. Use `async_inference` instead of `inference` and await calls like `await client.run(...)` and `await agent.send_message(...)`.
Can I build agents with custom tools?
Yes. Use the tool builder API to define client tools, app tools (calls to inference.sh apps), agent tools (delegation to sub-agents), and webhook tools (external APIs). Tools can require human approval before execution.
What models are available as core apps for agents?
Claude Sonnet 4, Claude 3.5 Haiku, GPT-4o, and GPT-4o Mini. Access via app references like `infsh/claude-sonnet-4@latest`.
Does the SDK handle file uploads?
Yes. Files can be auto-uploaded by passing file paths directly in input, or manually uploaded via `client.upload_file()` with optional metadata like custom filename and content-type.

Generated from the current SKILL.md. These answers refresh after source changes.