FoundationModels: On-Device LLM for iOS 26
Use Apple's FoundationModels framework to run language tasks on the device.
This skill covers:
- Text generation
- Typed output with
@Generable - Custom tools
- Live result streams
- Safe error handling
- Private and offline use
When to Use This Skill
Use this skill when an app needs to:
- Write or sum up text
- Turn plain text into Swift data
- Show an answer as it is made
- Let the model call app code
- Work without the cloud
- Keep user data on the device
Do not use it when:
- The device cannot run Apple Intelligence
- The task needs live web data and no local tool can provide it
- The input is too large for the model
- The task needs facts the on-device model may not know
Required Setup
Import the framework:
import FoundationModelsOnly use FoundationModels on supported OS and device versions. Add a fallback UI for older or unsupported devices.
Check Model Availability First
Check the model before you make a session:
struct GenerativeView: View {
private let model = SystemLanguageModel.default
var body: some View {
switch model.availability {
case .available:
ContentView()
case .unavailable(.deviceNotEligible):
Text("This device cannot use Apple Intelligence.")
case .unavailable(.appleIntelligenceNotEnabled):
Text("Turn on Apple Intelligence in Settings.")
case .unavailable(.modelNotReady):
Text("The language model is not ready yet.")
case .unavailable(let reason):
Text("The language model is not available: \(reason)")
}
}
}Availability can change. Check it again when the app returns to the front or before an important task.
Do not hide the whole app when the model is not ready. Keep other features working and show a clear fallback.
Basic Text Generation
Use a new session for a one-time task:
let session = LanguageModelSession()
do {
let response = try await session.respond(
to: "What is a good month to visit Paris?"
)
print(response.content)
} catch {
print("The request failed: \(error.localizedDescription)")
}Reuse one session when later prompts need the earlier chat:
let session = LanguageModelSession(instructions: """
You are a cooking helper.
Suggest meals from the food the user has.
Keep each answer short.
Do not claim that food is safe if you are not sure.
""")
let first = try await session.respond(
to: "I have chicken and rice."
)
let second = try await session.respond(
to: "Give me a meat-free choice instead."
)
print(first.content)
print(second.content)Good instructions should say:
- Who the model should act as
- What task it should do
- What shape the answer should have
- How long the answer should be
- What it must not do
- What to say when key facts are missing
Do not place secrets in instructions or prompts. On-device work helps privacy, but logs, crash reports, and tool code may still leak data.
Typed Output with @Generable
Use @Generable when your app needs fields, not free text.
Define a Type
@Generable(description: "Basic facts about a cat")
struct CatProfile {
var name: String
@Guide(description: "The cat's age", .range(0...20))
var age: Int
@Guide(description: "One short sentence about the cat's nature")
var profile: String
}Ask for That Type
let session = LanguageModelSession()
let response = try await session.respond(
to: "Make a profile for a cute rescue cat.",
generating: CatProfile.self
)
let cat = response.content
print("Name: \(cat.name)")
print("Age: \(cat.age)")
print("Profile: \(cat.profile)")Read results from response.content. Do not use response.output.
Guide Rules
Common guide rules include:
.range(0...20)for a number range.count(3)for an exact list sizedescription:for plain rules about a field
Keep each field clear and small. If a field may not be known, make it optional. Do not force the model to invent missing facts.
Check typed output before you save it or use it for an important action. A valid Swift value can still be wrong.
Custom Tool Calls
A tool lets the model ask your app to do a set task.
Define a Tool
struct RecipeSearchTool: Tool {
let name = "recipe_search"
let description = """
Find recipes that match a search term.
Return no more than the requested count.
"""
@Generable
struct Arguments {
var searchTerm: String
@Guide(description: "Number of recipes", .range(1...5))
var numberOfResults: Int
}
func call(arguments: Arguments) async throws -> ToolOutput {
let recipes = try await searchRecipes(
term: arguments.searchTerm,
limit: arguments.numberOfResults
)
let text = recipes
.map { "- \($0.name): \($0.description)" }
.joined(separator: "\n")
return .string(text.isEmpty ? "No recipes found." : text)
}
}Add the Tool to a Session
let session = LanguageModelSession(
tools: [RecipeSearchTool()]
)
let response = try await session.respond(
to: "Find three pasta recipes."
)
print(response.content)Handle Tool Errors
do {
let response = try await session.respond(
to: "Find a tomato soup recipe."
)
print(response.content)
} catch let error as LanguageModelSession.ToolCallError {
print("Tool failed: \(error.tool.name)")
if let recipeError = error.underlyingError as? RecipeSearchToolError {
switch recipeError {
case .databaseIsEmpty:
print("No recipes are stored.")
default:
print("Recipe search failed.")
}
}
} catch {
print("The request failed: \(error.localizedDescription)")
}Tool code must check all input. Treat model-made tool input like user input.
Ask the user before a tool does a risky action, such as:
- Deleting data
- Buying an item
- Sending a message
- Sharing private data
- Changing an account
Set time limits for slow tools. Support task canceling. Return short tool results so the session does not fill up.
Live Snapshot Streaming
A stream sends full partial values. It does not send text changes.
@Generable
struct TripIdeas {
@Guide(description: "Three short trip ideas", .count(3))
var ideas: [String]
}
let session = LanguageModelSession()
let stream = session.streamResponse(
to: "Give me three fun trip ideas.",
generating: TripIdeas.self
)
for try await partial in stream {
// Each field in PartiallyGenerated may be missing.
print(partial)
}Do not add each snapshot to the old one. Replace the old snapshot with the new snapshot.
SwiftUI Example
struct TripIdeasView: View {
let prompt: String
let session: LanguageModelSession
@State private var partialResult: TripIdeas.PartiallyGenerated?
@State private var errorMessage: String?
var body: some View {
List {
ForEach(partialResult?.ideas ?? [], id: \.self) { idea in
Text(idea)
}
}
.overlay {
if let errorMessage {
Text(errorMessage)
.foregroundStyle(.red)
.padding()
}
}
.task(id: prompt) {
partialResult = nil
errorMessage = nil
do {
let stream = session.streamResponse(
to: prompt,
generating: TripIdeas.self
)
for try await partial in stream {
try Task.checkCancellation()
partialResult = partial
}
} catch is CancellationError {
// The view closed or the prompt changed.
} catch {
errorMessage = "Could not make trip ideas."
}
}
}
}Update view state on the main actor when the code runs outside SwiftUI's task context.
One Full Example
This example turns a note into a small task list:
import FoundationModels
@Generable(description: "A task found in a user's note")
struct TaskItem {
var title: String
@Guide(description: "True only when the note gives a due date")
var hasDueDate: Bool
var dueDateText: String?
}
@Generable(description: "Tasks found in a user's note")
struct TaskList {
@Guide(description: "No more than five tasks")
var tasks: [TaskItem]
}
func readTasks(from note: String) async throws -> [TaskItem] {
let model = SystemLanguageModel.default
guard case .available = model.availability else {
return []
}
let cleanNote = note.trimmingCharacters(in: .whitespacesAndNewlines)
guard !cleanNote.isEmpty else {
return []
}
let session = LanguageModelSession(instructions: """
Find only tasks that the user clearly asked to do.
Do not make up dates.
Keep each title short.
Return no more than five tasks.
""")
let response = try await session.respond(
to: cleanNote,
generating: TaskList.self
)
return Array(response.content.tasks.prefix(5))
}Example input:
Call Sam tomorrow. Buy milk. Maya may visit next week.Expected kind of result:
Call Sam, due tomorrow
Buy milk, no due dateDo not turn “Maya may visit next week” into a task. It is not a clear action for the user.
Session Limits
A session handles one request at a time. Do not start two requests on the same session.
Before sending work, check:
guard !session.isResponding else {
return
}For work that may run at the same time, use separate sessions.
Keep instructions, prompts, tool results, chat history, and output small. They share the model's context space. If a request is too large:
- Split the input into small parts.
- Process one part at a time.
- Join the short results.
- Start a new session if old chat is no longer needed.
Do not split in the middle of a sentence or data record.
Error and State Rules
Handle these cases:
- The model is not available.
- The user turns off Apple Intelligence.
- The model is still being set up.
- The prompt is empty.
- The task is canceled.
- The app moves to the back.
- A tool fails or takes too long.
- The session is already busy.
- The input is too large.
- The model refuses the request.
- Typed output is valid but has bad facts.
Show simple error text to the user. Log only safe details. Never log a full private prompt by default.
Best Practices
- Check model availability before use.
- Give short and clear instructions.
- Use
@Generablefor data fields. - Read results from
response.content. - Replace old stream snapshots with new ones.
- Check model and tool output before important use.
- Keep prompts and tool results small.
- Use a new session when chat history is not needed.
- Cancel work when its screen closes.
- Give the user a useful fallback.
- Test on real supported devices.
- Use Xcode Instruments to check speed, memory, and power use.
Common Mistakes
Do not:
- Assume every device can use the model.
- send two requests through one session at once.
- Use
.outputinstead of.content. - Parse free text when a typed result fits.
- Trust made-up facts just because the type is valid.
- Put many steps into one large prompt.
- Keep old chat history forever.
- Append streamed snapshots as if they were changes.
- Let a tool act without checking its input.
- Let the model make risky changes without user approval.
- Claim that no data leaves the device if your tools, logs, or app code send it out.