All skills
inference-shell avatar

/python-sdk

@fbe0aa4

Python SDK for inference.sh - run AI apps, build agents, and integrate with all models. Package: inferencesh (pip install inferencesh). Supports sync/async, streaming, file uploads. Build agents with template or ad-hoc patterns, tool builder API, skills, and human approval. Use for: Python integration, AI apps, agent development, RAG pipelines, automation. Triggers: python sdk, inferencesh, pip install, python api, python client, async inference, python agent, tool builder python, programmatic ai, python integration, sdk python

Use this Skill: https://skilld.dev/gh/inference-shell/skills/python-sdk

This session only. Nothing lands on disk.

referencessessions.md

≈1.8k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Sessions Reference

Stateful execution with warm workers.

What Are Sessions?

Sessions keep workers warm between requests, enabling:

  • Faster execution - No cold start on subsequent calls
  • Shared state - Maintain context, loaded models, cached data
  • Cost efficiency - Reuse initialized resources

Creating a Session

from inferencesh import inference

client = inference(api_key="inf_...")

# Start new session
result = client.run({
    "app": "my-app",
    "input": {"action": "initialize"},
    "session": "new"
})

session_id = result["session_id"]
print(f"Session: {session_id}")

Using an Existing Session

# Continue in same session
result = client.run({
    "app": "my-app",
    "input": {"action": "process", "data": "..."},
    "session": session_id
})

Session Timeout

Set how long idle sessions stay alive (1-3600 seconds):

# 5-minute timeout
result = client.run({
    "app": "my-app",
    "input": {"action": "init"},
    "session": "new",
    "session_timeout": 300
})

Session Lifecycle

1. Create session (session: "new")
   ↓
2. Worker starts, initializes app
   ↓
3. Subsequent calls reuse worker (session: session_id)
   ↓
4. Idle timeout reached or explicit close
   ↓
5. Worker terminates

Use Cases

Model Loading

Load a model once, use it multiple times:

# Initial load (slow)
result = client.run({
    "app": "ml-inference",
    "input": {"action": "load_model", "model": "large-model-v2"},
    "session": "new",
    "session_timeout": 600
})
session_id = result["session_id"]

# Fast inference calls
for item in data_batch:
    result = client.run({
        "app": "ml-inference",
        "input": {"action": "predict", "data": item},
        "session": session_id
    })
    print(result["output"])

Browser Automation

Keep browser open across multiple actions:

# Start browser session
result = client.run({
    "app": "browser-automation",
    "input": {"action": "start", "url": "https://example.com"},
    "session": "new",
    "session_timeout": 300
})
session_id = result["session_id"]

# Navigate
client.run({
    "app": "browser-automation",
    "input": {"action": "click", "selector": "#login-btn"},
    "session": session_id
})

# Fill form
client.run({
    "app": "browser-automation",
    "input": {"action": "type", "selector": "#username", "text": "user@example.com"},
    "session": session_id
})

# Take screenshot
result = client.run({
    "app": "browser-automation",
    "input": {"action": "screenshot"},
    "session": session_id
})

Stateful Conversations

# Initialize chat context
result = client.run({
    "app": "chat-with-memory",
    "input": {"action": "init", "system": "You are a helpful assistant."},
    "session": "new",
    "session_timeout": 1800  # 30 minutes
})
session_id = result["session_id"]

# Multi-turn conversation
messages = [
    "What is quantum computing?",
    "Can you give me a simple example?",
    "How is it different from classical computing?"
]

for msg in messages:
    result = client.run({
        "app": "chat-with-memory",
        "input": {"message": msg},
        "session": session_id
    })
    print(f"Assistant: {result['output']['response']}")

Data Processing Pipeline

# Load data once
result = client.run({
    "app": "data-processor",
    "input": {"action": "load", "dataset": "large_dataset.parquet"},
    "session": "new",
    "session_timeout": 900
})
session_id = result["session_id"]

# Run multiple analyses
analyses = ["summary", "correlations", "outliers", "trends"]

for analysis in analyses:
    result = client.run({
        "app": "data-processor",
        "input": {"action": "analyze", "type": analysis},
        "session": session_id
    })
    print(f"{analysis}: {result['output']}")

Session Management

Check Session Status

# Sessions are implicitly active when used
# If session expired, you'll get an error
try:
    result = client.run({
        "app": "my-app",
        "input": {"action": "check"},
        "session": session_id
    })
except Exception as e:
    if "session not found" in str(e).lower():
        print("Session expired, creating new one")
        # Create new session

Explicit Session Close

# Close session to free resources
client.run({
    "app": "my-app",
    "input": {"action": "cleanup"},
    "session": session_id
})
# Session will terminate after this call

Session Recovery Pattern

class SessionManager:
    def __init__(self, client, app, timeout=300):
        self.client = client
        self.app = app
        self.timeout = timeout
        self.session_id = None

    def ensure_session(self):
        if self.session_id is None:
            result = self.client.run({
                "app": self.app,
                "input": {"action": "init"},
                "session": "new",
                "session_timeout": self.timeout
            })
            self.session_id = result["session_id"]
        return self.session_id

    def run(self, input_data):
        try:
            return self.client.run({
                "app": self.app,
                "input": input_data,
                "session": self.ensure_session()
            })
        except Exception as e:
            if "session" in str(e).lower():
                # Session expired, create new one
                self.session_id = None
                return self.client.run({
                    "app": self.app,
                    "input": input_data,
                    "session": self.ensure_session()
                })
            raise

# Usage
manager = SessionManager(client, "my-app", timeout=600)
result = manager.run({"action": "process", "data": "..."})

Async Sessions

from inferencesh import async_inference
import asyncio

async def session_workflow():
    client = async_inference(api_key="inf_...")

    # Create session
    result = await client.run({
        "app": "my-app",
        "input": {"action": "init"},
        "session": "new",
        "session_timeout": 300
    })
    session_id = result["session_id"]

    # Run operations
    tasks = [
        client.run({
            "app": "my-app",
            "input": {"action": "process", "id": i},
            "session": session_id
        })
        for i in range(10)
    ]

    # Note: These run sequentially on the same worker
    results = []
    for task in tasks:
        results.append(await task)

    return results

asyncio.run(session_workflow())

Best Practices

  1. Set appropriate timeouts - Balance between keeping workers warm and resource usage
  2. Handle session expiry - Always catch and handle session not found errors
  3. Clean up when done - Close sessions explicitly if you know you're finished
  4. Don't over-parallelize - Session requests go to the same worker sequentially
  5. Monitor costs - Long-running sessions incur ongoing charges

Source: SKILL.md on GitHub

3 warnings16d5 checks · Risk MEDIUM
  • Gen Agent Trust Hub16d

    The skill provides a Python SDK for interacting with the inference.sh platform. While intended for development, it contains documentation examples that promote insecure coding practices, such as using `eval()` on unsanitized tool arguments. It also features automatic file upload capabilities that could be exploited to exfiltrate sensitive local files if an agent is tricked into processing malicious paths.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer6mo

    1/7 files flagged

  • ZeroLeaks5mo

    3 findings · Score: 69/100

Signed by skilld at fbe0aa4. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 3 months ago
What it can do
Runs commands
All 2 allowed tools
Bash(pip install inferencesh)Bash(python *)
  • Python
  • sdk
  • inference-sh
  • ai-agents
  • llm
  • async
  • streaming
  • tool-builder
  • rag
  • api-client

README badge

README badge for inference-shell/skills/python-sdk

Provides a Python SDK for inference.sh that runs AI apps, builds agents with tool builders and human approval workflows, and integrates 250+ models via a single API. Supports sync/async execution, streaming, file uploads, and stateful sessions—targets Python developers building AI applications and agent systems.

Generated from the current SKILL.md.

What Python versions does this SDK support?
Python 3.8 and later. Install with `pip install inferencesh` for sync support or `pip install inferencesh[async]` for async/await.
Does this support async/await?
Yes. Use `async_inference` instead of `inference` and await calls like `await client.run(...)` and `await agent.send_message(...)`.
Can I build agents with custom tools?
Yes. Use the tool builder API to define client tools, app tools (calls to inference.sh apps), agent tools (delegation to sub-agents), and webhook tools (external APIs). Tools can require human approval before execution.
What models are available as core apps for agents?
Claude Sonnet 4, Claude 3.5 Haiku, GPT-4o, and GPT-4o Mini. Access via app references like `infsh/claude-sonnet-4@latest`.
Does the SDK handle file uploads?
Yes. Files can be auto-uploaded by passing file paths directly in input, or manually uploaded via `client.upload_file()` with optional metadata like custom filename and content-type.

Generated from the current SKILL.md. These answers refresh after source changes.