Skip to main content
First-party Python SDK for the Interhuman API. Quickstart:

Clients & helpers

AuthClient

Client for the token endpoints (/v1/auth and /v1/client_tokens). Args: environment: Named environment to call. Defaults to production. base_url: Explicit base URL override (e.g. http://localhost:8080). http_client: Optional shared httpx.AsyncClient. When omitted, a client is created per request. Constructor

AuthClient.base_url

The resolved HTTP base URL this client calls.

AuthClient.create_client_token()

Mint a short-lived, capped client token for direct end-user use. Call this from a trusted server: the full API key travels in the request body. Omitted options use the server defaults (scope interhumanai.stream, 300 second lifetime clamped to 60-3600). Args: api_key: The full API key (ih_...) minting the token. scopes: Scopes the client token should carry. expires_in: Requested lifetime in seconds (server clamps to 60-3600). max_duration_seconds: Cap on a single live session’s duration. max_bytes: Cap on bytes accepted across the token’s sessions. max_concurrent: Cap on concurrent sessions (server default 1). max_video_seconds: Video-seconds budget across all surfaces. allowed_origins: Browser origins allowed to use the token. Returns: The minted client token and its effective caps.

AuthClient.create_token()

Exchange API key credentials for a short-lived bearer access token. Args: key_id: The API key id. key_secret: The API key secret. scopes: Scopes to request; must be non-empty and held by the key. Returns: The minted token, its lifetime, and the granted scopes.

AuthClient.revoke_client_token()

Revoke a previously minted client token. Revoking an already-expired token is a no-op and also succeeds. Args: api_key: The API key that minted the token. token: The client token to revoke.

InterhumanClient

High-level client for the Interhuman API. Authenticate either with API key credentials (key_id + key_secret, exchanged for short-lived bearer tokens that are refreshed automatically) or with a pre-issued access_token used as-is. Exactly one of the two must be provided. The client exposes every public surface: auth for token endpoints, upload for complete files, and stream / realtime for live WebSocket sessions. Args: key_id: API key id (paired with key_secret). key_secret: API key secret. access_token: Pre-issued bearer token (JWT or client token). scopes: Scopes requested when exchanging credentials. Defaults to upload + stream. Ignored when access_token is used. environment: Named environment to call. Defaults to production. base_url: Explicit base URL override (e.g. http://localhost:8080). http_client: Optional shared httpx.AsyncClient for the HTTP surfaces. The caller owns its lifecycle. refresh_skew_seconds: How long before expiry managed tokens refresh. Constructor

InterhumanClient.get_token()

Return the bearer token the client is currently using. With managed credentials this mints or refreshes as needed; with a pre-issued access_token it returns that token unchanged.

InterhumanClient.realtime()

Create a client for one WS /v0/realtime/analyze session. Each call returns a fresh, unconnected client; a client handles a single session.

InterhumanClient.stream()

Create a client for one live stream session. Each call returns a fresh, unconnected client; a client handles a single session. api_version="v1" (the default) opens WS /v1/stream/analyze, analyzed by Inter-1; api_version="v2" opens WS /v2/stream/analyze, analyzed by Inter-2. The protocol and the events are identical on both. Args: api_version: Which stream endpoint to open, "v1" or "v2". model: The Inter-2 model a "v2" session opens on (inter-2 when omitted). The credential must carry interhumanai.stream.<model> for it; see ~interhumanai.StreamClient.

RealtimeClient

One live session against WS /v0/realtime/analyze. The realtime endpoint ships under the v0 path and requires the interhumanai.realtime scope. It offers multi-track analysis configuration, caller-supplied transcripts, and periodic recommendations.

RealtimeClient.send_transcript()

Send the latest client-side transcript of the conversation. Each call fully replaces any previously sent transcript; the latest transcript is rendered into subsequent recommendation output. Args: transcript: Ordered transcript segments (speaker ids are zero-based).

RealtimeClient.update_config()

Replace the session configuration. Each update fully replaces the previous configuration; omitted options reset to their server defaults. The server acknowledges with a session.updated event. The recommendation step’s system prompt and model are managed by Interhuman and cannot be set per session. Args: analysis_groups: Analysis tracks to run (must be non-empty when provided). Server default when omitted: audio and visual. The visual group reads the picture, so pass ["audio"] alone to analyze a stream whose recording carries no video track: leaving visual selected and sending such a stream stops the session’s analysis with a single ih5001 error, the audio tracks included. Narrowing to ["audio"] resumes analysis from the next window. realtime_recommendation_frequency: How often recommendation output is generated. realtime_recommendation_instructions: Non-empty instructions that enable recommendations; when omitted or empty, recommendations are disabled. goal_dimensions: Goal dimensions that enable recommendation generation.

SessionSocketClient

One live analysis session over a WebSocket. Connect with connect (or async with), send binary video chunks with send_video, and consume typed server events by iterating the client with async for. Iteration ends when the connection closes; close_info then holds the close code and reason. Args: token_provider: Source of bearer tokens for authentication. environment: Named environment to connect to. Defaults to production. base_url: Explicit HTTP base URL override; converted to ws(s)://. Constructor

SessionSocketClient.close()

Close the WebSocket immediately, discarding in-flight analysis. Args: code: WebSocket close code to send. reason: Optional close reason.

SessionSocketClient.close_info

Close code and reason once the connection has ended, else None.

SessionSocketClient.connect()

Open the WebSocket connection and start receiving events. Returns: This client, for chaining. Raises: InterhumanConfigError: If the client was already connected. InterhumanError: If the handshake fails (bad credentials or scope, unreachable host).

SessionSocketClient.is_open

Whether the connection is currently open.

SessionSocketClient.request_close()

Ask the server to drain in-flight analysis and end the session. The server acknowledges with session.closing, emits any final events, sends session.ended, and closes the socket normally. Keep iterating to observe the drain; for an immediate teardown use close instead.

SessionSocketClient.send_video()

Send one binary video chunk. The first chunk must carry the container’s init header; later chunks are continuation fragments of the same WebM or fragmented-MP4 stream. The SDK does not enforce a chunk size; the server rejects chunks above its limit (32 MB by default, reported in session.ready as max_segment_size_bytes) with an ih6002 error. Args: chunk: The raw video bytes to send.

SessionSocketClient.url

The WebSocket URL this client connects to. Carries the endpoint’s handshake parameters and the SDK attribution query pair alongside any query parameters the configured base URL already had. The same identity also travels in the handshake’s X-Interhuman-SDK header; the two always agree.

SessionSocketClient.wait_closed()

Wait until the connection has fully closed. Returns: The session’s close code and reason.

SessionSocketClient.wait_for_session_ready()

Wait for the server’s session.ready event. The event is also delivered through iteration; this helper simply awaits it (or returns it if it already arrived). Returns: The session.ready event. Raises: InterhumanError: If the connection closes before the session becomes ready. The message carries the server’s own error envelope when one arrived before the close — a session the server accepts and then refuses, such as ih1003 — and the scope hint otherwise, which is the usual cause when the close explains nothing itself.

StaticTokenProvider

Token provider that always returns the same pre-issued token. Constructor

StaticTokenProvider.get_token()

Return the configured token.

StreamClient

One live session against WS /v1/stream/analyze or WS /v2/stream/analyze. Send WebM or fragmented-MP4 chunks with send_video; consume typed events with async for. api_version selects the endpoint: "v1" (the default) is analyzed by the Inter-1 model and requires the interhumanai.stream scope; "v2" is analyzed by an Inter-2 model. The session protocol and the events are identical on both, and each names the endpoint it actually opened in the errors it raises: a v2 session reports the Inter-2 stream, a v1 session the Stream. On "v2" the model is a permission of its own. model names the model the session opens on (StreamModel.INTER_2 when omitted), sent as the handshake’s model query parameter, and the credential must carry interhumanai.stream.<model> for it or the server refuses the session with ih2003. session.ready lists, under supported_session_config_options.model, the models the deployment serves that the credential may select; a later update_config(model=...) switching to one the credential lacks is answered with a non-fatal ErrorEvent (ih2003) while the session continues on its current model. Adding an audio_format on StreamModel.INTER_2_AUDIO lets you send raw PCM frames with send_audio instead of container media. Args: token_provider: Source of bearer tokens for authentication. environment: Named environment to connect to. Defaults to production. base_url: Explicit HTTP base URL override; converted to ws(s)://. api_version: Which stream endpoint to open, "v1" or "v2". model: The Inter-2 model a "v2" session opens on. Rejected on "v1", which offers no selection, and for StreamModel.INTER_2_DEEP, which is an upload model. Constructor

StreamClient.api_version

The stream endpoint this client opens ("v1" or "v2").

StreamClient.model

The model a "v2" session opens on, as named at construction. None on "v1" and when the server default (inter-2) is left to apply.

StreamClient.send_audio()

Send one frame of raw PCM audio. Only meaningful on api_version="v2" after update_config declared an audio_format with StreamModel.INTER_2_AUDIO; the server then reads every binary frame as signed 16-bit mono PCM at the declared sample rate. Frames can be any size and cut at any byte boundary. Without that declaration a binary frame is container media, which send_video names. Args: frame: The raw sample bytes to send.

StreamClient.update_config()

Replace the session configuration. Each update fully replaces the previous include and goal_dimensions; omitted options reset to their server defaults. model and audio_format are session state instead: omitting them keeps the values in force, and a different value is accepted only before the first media frame. This method never sends an explicit null for either, so a selection cannot be reset through it; open a new session instead. The server acknowledges with a session.updated event. Args: include: Optional sections to include in quality updates. goal_dimensions: Goal dimensions that enable feedback generation. model: The Inter-2 model to analyze the session with, on api_version="v2" only. StreamModel.INTER_2 (the default) reads the video; StreamModel.INTER_2_AUDIO hears the audio alone. audio_format: Declares that binary frames are raw PCM audio (see ~interhumanai.RawAudioFormat), sent with send_audio. Requires StreamModel.INTER_2_AUDIO.

TokenManager

Mints bearer tokens from API key credentials and refreshes them early. Tokens are cached until expires_in - refresh_skew_seconds elapses; concurrent callers share a single in-flight mint. Args: auth_client: The AuthClient used to mint tokens. key_id: The API key id. key_secret: The API key secret. scopes: Scopes to request on every mint. refresh_skew_seconds: Seconds before expiry at which the cached token is considered stale. clock: Monotonic clock returning seconds; injectable for tests. Constructor

TokenManager.get_token()

Return a valid access token, minting or refreshing when needed.

TokenManager.invalidate()

Drop the cached token so the next call mints a fresh one.

TokenProvider

Anything that can produce a bearer access token on demand. Constructor

TokenProvider.get_token()

Return a currently valid bearer access token.

UploadClient

Client for analyzing complete files. analyze is the v1 route: the file is analyzed inside the request and the report comes back with the response. submit, get_job and wait_for_job are the v2 job routes: a file is accepted as a job, analyzed by an Inter-2 model after the response, and read back by id until it is completed or failed. Args: token_provider: Source of bearer tokens for authentication. environment: Named environment to call. Defaults to production. base_url: Explicit base URL override. http_client: Optional shared httpx.AsyncClient. When omitted, a client is created per request. Constructor

UploadClient.analyze()

Analyze a complete video file and return the detected signals. The video must be at least 3 seconds long and at most 32 MB, in one of the supported containers (mp4, avi, mov, mkv, mpeg-ts, webm). Args: file: The video to analyze - raw bytes, an open binary file object, or a filesystem path. filename: Filename reported to the API. Defaults to the path or file object’s name, else video. content_type: MIME type of the file (e.g. video/mp4). include: Optional response sections to include (conversation quality overall and/or timeline). goal_dimensions: Goal dimensions that trigger interaction feedback. Ignored when conversation_context is set. conversation_context: Free-text description of the interaction; when set, it drives feedback generation. Returns: The analysis result. Optional sections are only present when requested.

UploadClient.get_job()

Read a job’s current envelope from GET /v2/upload/jobs/{job_id}. Args: job_id: The job’s id, or the envelope submit returned. Returns: The envelope as it stands: result once completed, error once failed. Raises: InterhumanAPIError: ih4021 (404) when the id is unknown to this account or the job has expired, among the usual errors.

UploadClient.submit()

Submit a file to POST /v2/upload/analyze as an asynchronous job. The API validates the file and answers with the job envelope before the analysis runs. With wait_seconds above zero the request is held open for up to that long, and the envelope comes back terminal (result or error set) when the job finished in time; otherwise it comes back queued or running and wait_for_job or get_job reads it later. For UploadModel.INTER_2_AUDIO the file is wav, flac, mp3, m4a, ogg, or a webm or mp4 with an audio track — at least 3 seconds of media, at most 32 MB, and no longer than the deployment’s maximum duration (30 minutes by default; a longer file raises ih4004). The job cuts it into fixed windows and analyzes each one. Args: file: The file to analyze - raw bytes, an open binary file object, or a filesystem path. model: The Inter-2 model to analyze with. INTER_2 and INTER_2_DEEP are valid values the route does not serve yet and raise ~interhumanai.InterhumanAPIError (ih4020). wait_seconds: How long the API may hold the request waiting for the job to finish. 0 (the default) answers at once. The deployment bounds it; a larger value raises ih4005. filename: Filename reported to the API. Defaults to the path or file object’s name, else video. content_type: MIME type of the file (e.g. audio/wav). Returns: The job envelope. Check UploadJob.status.

UploadClient.wait_for_job()

Poll a job until it is completed or failed, and return it. A terminal envelope passed in is returned at once without a request. A failed job is returned, not raised: read UploadJob.error. Args: job_id: The job’s id, or an envelope from submit or get_job. timeout: Give up after this many seconds. None waits until the job is terminal. poll_interval: Seconds between status reads. Returns: The terminal envelope. Raises: UploadJobTimeoutError: timeout elapsed first. The job keeps running; the error carries the last envelope read.

Data models

AnalysisResult

Response of POST /v1/upload/analyze. Optional sections are only present when requested: feedback when goal dimensions or a conversation context were supplied, conversation_quality when requested via include flags.

ClientTokenResponse

Response of POST /v1/client_tokens.

CloseInfo

Close code and reason of a finished WebSocket session.

ConversationQuality

Conversation-quality section of an analysis result.

ConversationQualityTimelineEntry

Conversation-quality scores for one slice of the video.

ConversationQualityUpdatedData

Latest conversation-quality scores (sections follow the include flags).

ConversationQualityUpdatedEvent

New conversation-quality scores are available.

ConversationQualityValues

Conversation-quality scores (0-100; 50 means no evidence either way).

CoverageDegradedData

Ranges analyzed with partial visual coverage (billed normally). Unlike coverage.dropped, these windows were analyzed — the audio and any decodable video informed the analysis; only the visual coverage was partial (e.g. a screen share whose keyframe interval exceeds the analysis window).

CoverageDegradedEvent

One or more windows were analyzed with partial visual coverage.

CoverageDroppedData

Time ranges no analysis covers (not billed). Emitted when the analysis pipeline sheds buffered video under backpressure, and when the incoming stream itself skipped ahead — a client stall whose media never arrived while its recorder clock kept running. Either way the listed ranges were never analyzed.

CoverageDroppedEvent

Some video was dropped without being analyzed.

CoverageRange

A time range of video that was not analyzed.

EngagementStateEntry

An engagement state over a time range.

EngagementUpdatedData

The engagement state entered at start.

EngagementUpdatedEvent

The subject’s engagement state changed.

ErrorData

A session error notice (fatal errors also close the socket).

ErrorEvent

The server reported an error for this session.

Feedback

Actionable interaction feedback generated from the analysis.

FeedbackGeneratedData

Feedback text generated from the active goal dimensions.

FeedbackGeneratedEvent

Interaction feedback was generated.

NoFeedback

Explicit indication that no feedback was warranted, with the reason.

RawAudioFormat

Declares that a session’s binary frames are raw audio, not container media. Signed 16-bit little-endian PCM at sample_rate, mono, with no header; frames may be cut at any byte boundary. Only inter-2-audio accepts it. Attributes: encoding: Always pcm_s16le. sample_rate: Samples per second, 16000 or 24000. channels: Always 1.

RealtimeRecommendationGeneratedData

Generated guidance for the analyzed window. text is the literal NO_GUIDANCE when the model had nothing to suggest for this window.

RealtimeRecommendationGeneratedEvent

New recommendation output is available.

RealtimeSessionConfigOptions

Session-config options the realtime endpoint supports. The recommendation step’s system prompt and model are managed by Interhuman and are not session-selectable, so only the caller-owned options appear here.

RealtimeSessionReadyData

Limits and supported options of a newly opened realtime session.

RealtimeSessionReadyEvent

The session is accepted and ready to receive video.

RealtimeSessionUpdatedData

The realtime session configuration now in effect. Reports only the caller-owned options. The recommendation system prompt and model are managed by Interhuman rather than session state, so they are not part of the acknowledgment.

RealtimeSessionUpdatedEvent

Acknowledgment of a session-config update.

RealtimeSignalDetectedData

A newly detected signal (its end is not known yet). The realtime payload carries no rationale. modality names the analyses whose evidence produced the signal: one or both of "audio" and "visual".

RealtimeSignalDetectedEvent

A social signal was detected.

RealtimeSignalUpdatedData

Updated details for a signal that is still active. The realtime payload carries no rationale. A change to modality — the set of analyses reporting the signal — is one of the things that emits signal.updated.

RealtimeSignalUpdatedEvent

An active signal’s details changed.

SessionClosingData

Drain window granted after a graceful close request.

SessionClosingEvent

The server accepted a graceful close and is draining.

SessionConfigOptions

Session-config options the stream endpoint supports. Attributes: include: The conversation-quality sections a config may include. goal_dimensions: The goal dimensions a config may set, or None when feedback is disabled. model: The models this deployment serves on WS /v2/stream/analyze; None on WS /v1/stream/analyze, which offers no selection.

SessionEndedData

Why the session ended.

SessionEndedEvent

Final envelope of a gracefully ended session.

SessionReadyData

Limits and supported options of a newly opened stream session.

SessionReadyEvent

The session is accepted and ready to receive video.

SessionUpdatedData

The session configuration now in effect. Attributes: include: The conversation-quality sections in force. goal_dimensions: The goal dimensions in force, or None when feedback is disabled. model: The model analyzing the session on WS /v2/stream/analyze; None on WS /v1/stream/analyze.

SessionUpdatedEvent

Acknowledgment of a session-config update.

Signal

A detected social signal over a time range. modality names the analysis modalities that detected this signal. When several tracks detect the same signal, every contributing modality is included.

SignalDetectedData

A newly detected signal (its end is not known yet). modality names the analyses whose evidence produced the signal. For stream sessions this is ["video"].

SignalDetectedEvent

A social signal was detected.

SignalEndedData

End time of a signal that is no longer active.

SignalEndedEvent

An active signal ended.

SignalUpdatedData

Updated details for a signal that is still active. A change to modality — the set of analyses reporting the signal — is one of the things that emits signal.updated.

SignalUpdatedEvent

An active signal’s details changed.

TokenResponse

Response of POST /v1/auth.

TranscriptSegment

One segment of a conversation transcript.

UnknownEvent

A server envelope whose type this SDK version does not know. Newer API versions may add event types; they are surfaced as-is instead of failing the session. Attributes: type: The envelope’s type discriminator. raw: The full envelope payload as received.

UploadJob

The job envelope both v2 upload routes answer with. Returned by POST /v2/upload/analyze and GET /v2/upload/jobs/{job_id} alike. result is set only when status is COMPLETED; error only when it is FAILED. status_url is the status resource’s path, relative to the API base URL.

UploadJobError

Why an upload job failed, in the API’s standard error-body shape.

UploadJobResult

The result of a completed upload job: one entry per analyzed window.

UploadJobWindow

The analysis of one fixed-length window of an upload job’s file. signals carry the window’s span as their start and end, and their modality names the evidence the model read (["audio"] for inter-2-audio).

Enumerations

AnalysisGroup

Analysis track groups selectable on the realtime API.

EngagementLevel

Coarse engagement state of the analyzed subject.

GoalDimension

Interaction-goal dimensions used for feedback and conversation quality.

IncludeFlag

Optional response sections selectable on upload and stream analysis.

Probability

Confidence band attached to a detected signal.

RealtimeRecommendationFrequency

How often the realtime API generates feedback output. HIGH is roughly every 10 seconds of analyzed video, MEDIUM every 20 seconds, and LOW every 30 seconds.

Scope

OAuth-style scopes accepted by the Interhuman API. A scope names one operation, or one (operation, model) cell of the Inter-2 routes: the bare UPLOAD and STREAM are the Inter-1 cells (POST /v1/upload/analyze, WS /v1/stream/analyze), and each UPLOAD_INTER_2* / STREAM_INTER_2* member is one model on the matching /v2 route. Any UPLOAD_INTER_2* member also allows reading the account’s jobs at GET /v2/upload/jobs/{job_id}. No scope implies another, and there is no stream scope for inter-2-deep: it is an upload-only model.

SignalType

Social signals the Interhuman API can detect. TENSION and FRUSTRATION name the same underlying negative state, reported separately so the detecting source stays legible: TENSION comes from the Realtime API’s visual track.

StreamModel

Inter-2 models a WS /v2/stream/analyze session can select. INTER_2 (the default) reads the video. INTER_2_AUDIO hears the audio alone: an audio-only recording is accepted, the video of a recording that carries one is dropped, and raw PCM frames are accepted once a ~interhumanai.RawAudioFormat is declared. INTER_2_DEEP is an upload model; naming it on the stream is refused. WS /v1/stream/analyze offers no selection and rejects any value.

UploadJobStatus

Lifecycle of a POST /v2/upload/analyze job. QUEUED and RUNNING are transient; COMPLETED and FAILED are terminal, and the job stays readable until its expires_at.

UploadModel

Inter-2 models a POST /v2/upload/analyze job can name. INTER_2_AUDIO is served today. INTER_2 and INTER_2_DEEP are valid values the route does not serve yet; submitting them raises an ~interhumanai.InterhumanAPIError with error_id ih4020.

Exceptions

InterhumanAPIError

An error response from the Interhuman API, or a transport failure. Attributes: status: HTTP status code, or 0 when the request never reached the API (network failure). error_id: Machine-readable error code from the response body, when the API supplied one. correlation_id: Correlation id of the failed request, when supplied. link: Documentation link for the error, when supplied. body: The parsed JSON error body, when one was returned.

InterhumanConfigError

Raised for client-side misuse, before any network call is made. Examples: missing credentials, sending on a socket that is not open, or connecting a client that is already connected.

InterhumanError

Base class for every error raised by the Interhuman SDK.

UploadJobTimeoutError

An upload job did not reach a terminal state within the wait the caller allowed. Raised by ~interhumanai.UploadClient.wait_for_job. The job itself is unaffected — it keeps running, and job is the last envelope read, so the caller can keep polling with its job_id.

Functions

http_to_ws_base_url()

Derive the WebSocket base URL from an HTTP base URL. https:// becomes wss:// and http:// becomes ws://; any path is preserved and a trailing slash is trimmed. Args: http_base_url: The HTTP base URL to convert. Returns: The WebSocket base URL without a trailing slash.

parse_realtime_event()

Parse one realtime envelope into its typed event model. Args: payload: The decoded JSON envelope (must carry a string type). Returns: The matching typed event, or UnknownEvent for a type this SDK version does not know.

parse_stream_event()

Parse one stream envelope into its typed event model. Args: payload: The decoded JSON envelope (must carry a string type). Returns: The matching typed event, or UnknownEvent for a type this SDK version does not know.

resolve_http_base_url()

Resolve the HTTP base URL from an explicit override or a named environment. Args: base_url: Explicit base URL (e.g. http://localhost:8080). Takes precedence over environment when provided. environment: Named environment. Defaults to production. Returns: The base URL without a trailing slash.

sdk_header_value()

Return the X-Interhuman-SDK value: python/<version>.

Constants

DEFAULT_JOB_POLL_INTERVAL_SECONDS

Type: float

DEFAULT_REFRESH_SKEW_SECONDS

Type: float

DEFAULT_SCOPES

Type: tuple

Environment

Type: type alias

HTTP_BASE_URLS

Type: dict

InteractionFeedback

Type: type alias

NO_GUIDANCE

Type: str

PcmSampleRate

Type: type alias

RealtimeEvent

Type: type alias

SDK_HEADER_NAME

Type: str

SDK_NAME

Type: str

STREAM_ENDPOINT_PATHS

Type: dict

StreamApiVersion

Type: type alias

StreamEvent

Type: type alias

VideoInput

Type: type alias