Picsha AI

Search Infrastructure

Digital Asset Managers often struggle with metadata. Expecting humans to painstakingly tag every uploaded image is an archaic approach.

The moment a media file is ingested, it is enqueued to Picsha's asynchronous processing workers, which automatically extract structural analysis from the image via AWS Rekognition and visual embeddings via Amazon Titan Multimodal, streaming the result into our underlying OpenSearch Vector Database.

Standard Search (Keyword)

By default, triggering a Standard Search executes an ultra-fast query seeking direct matches against the original_name, extracted text (OCR), and any manual metadata tags attached.

This yields instantaneous results but relies heavily on precise keyword matching.

AI Search (Hybrid Mode)

Passing mode: "ai" to our Search endpoint unleashes our most powerful API abstraction.

When you query our servers with conversational NLP (e.g. "Show me the photos of dogs running on the beach from last Tuesday"):

  1. Agentic Parsing: We route the string through an LLM layer that parses actionable metadata constraints (e.g., isolating last Tuesday into a literal timestamp boundary).
  2. Text-to-Image Embedding: The rest of the semantic request ("dogs running on the beach") is converted into a high-dimensional multimodal vector array.
  3. k-NN Vector Execution: We perform a blistering k-NN (k-nearest neighbors) query against your Organization's dedicated vector space in OpenSearch.

Strict Intent Filtering (Methodology)

To prevent multimodal embeddings from surfacing false positives (e.g., searching for "Bill laughing" and getting dozens of unrelated but highly scored laughing people), our Agentic Parsing layer enforces strict, absolute metadata constraints prior to executing the vector search:

  • Visual Concepts & Colors: Handled purely by vector distances. If you search for "pink flowers", no explicit filters are applied. The embedding inherently calculates and surfaces visually matching images.
  • Human Names (Entities): Treated as non-negotiable hard filters. If the parsing layer detects a human name (e.g., "Hello bill"), it interprets the goal as a definitive entity search. It generates an explicit OpenSearch filter demanding that the asset must possess a match inside the ai.faces or user.tags indices.
  • Methodology Impact: While this guarantees exceptionally high accuracy for photographic face retrieval, it means text-heavy documents containing a person's name buried in a paragraph summary (like an ai.description) will be intentionally dropped from AI search pools unless the asset is explicitly tagged.

Search Thresholds

AI Search returns hits ordered by relevance score (matchScore). To prevent hallucinated matches from muddying the results, our backend inherently filters out any assets falling below a relevance threshold.

When building applications, you can dictate this exact cutoff threshold explicitly:

{
  "query": "corporate headshots with blue backgrounds",
  "mode": "ai",
  "threshold": 0.65 
}

Advanced Filters

In addition to semantic queries, you can supply structured filters in the advancedFilters (or filters) block to limit search scope. These filters are applied prior to vector ranking:

  • Date Ranges:
    • addedAfter / addedBefore: Boundaries for the ingestion timestamp (e.g. "now-30d").
    • capturedAfter / capturedBefore: Boundaries for the EXIF photo capture timestamp.
  • AI Metadata:
    • face: Match a recognized face identity (e.g. "John Doe").
    • object: Match an auto-detected item label (e.g. "dog").
    • place: Match reverse-geocoded location strings (e.g. "Denver").
    • ocr: Match text found within images or documents.
  • Media Context:
    • type: Match the MIME type prefix (e.g. "image" or "video"), or a list of prefixes (["image/png", "image/jpeg"]) any of which may match.
    • style: Match custom aesthetic or UI categories (e.g. "STANDARD", "FILM", "URL").
  • Custom Metadata:
    • metadataField / metadataValue: Search a specific metadata field (e.g., legacy fields like "pictures_meta_subject.value", or custom fields in user.custom).
    • minRating: Integer 1–5. Only assets whose metadata.rating is at least this many stars (so 3 means three stars or better).
    • metadata: Object of exact matches on custom metadata keys, e.g. {"asset_status": "Retired"}. Every key must match. Unlike metadataField, several keys can be combined. A key may list several values ({"asset_visibility": ["@all", "@sales"]}), any of which matches; on a key that stores an array, any stored element may match any listed value (up to 200 values per key).
  • Alternatives:
    • anyOf: A list of up to 10 filter blocks, each with the same fields as advancedFilters itself (minus anyOf). An asset matches when it satisfies at least one block; the whole list is combined with the other filters. See the example below.

Example Filter Request

{
  "query": "corporate headshots",
  "mode": "ai",
  "advancedFilters": {
    "type": "image",
    "capturedAfter": "2026-01-01",
    "style": "STANDARD",
    "metadataField": "pictures_meta_subject.value",
    "metadataValue": "#special events"
  }
}

Example: alternatives with anyOf

{
  "query": "beach",
  "advancedFilters": {
    "type": "image",
    "anyOf": [
      { "metadata": { "asset_visibility": ["@sales"] } },
      { "metadata": { "asset_visibility": ["@all"], "approval": "approved" } }
    ]
  }
}

Images matching "beach" that are either visible to @sales, or visible to @all and approved. Inside a block the filters are ANDed; blocks are ORed with each other. The same metadata and anyOf filters are available on GET /v1/assets as JSON query parameters, so a listing and a search can apply one policy.

Scope

Search covers the same assets as GET /v1/assets: an organization-scoped API key searches the whole organization, and a personal key searches its own uploads. Trashed assets are never returned.

Facets (Distinct Values with Counts)

Add a facets array to any search request to get the distinct values of a field, with counts, over the same scope as the hits — the caller's organization, the text query (standard mode), and any advancedFilters. Use it to build filter sidebars, people/place pickers, tag clouds, or autocomplete without scanning your library.

POST /search?limit=0
{
  "query": "",
  "mode": "standard",
  "advancedFilters": { "capturedAfter": "now-30d/d" },
  "facets": [
    { "field": "faces",  "size": 1000, "order": "key" },
    { "field": "places", "size": 1000, "order": "key" },
    { "field": "tags" }
  ]
}
FieldTypeNotes
fieldenumfaces, labels, places, tags, mimeType, cameraMake, cameraModel
sizeint, 1–1000, default 50Maximum number of values returned
ordercount | keycount (default) = most frequent first; key = alphabetical

Up to 10 facets per request. Pass ?limit=0 to receive facets and the total without any hits.

{
  "results": [],
  "pagination": { "page": 1, "limit": 0, "total": 4312 },
  "facets": {
    "faces":  [{ "value": "Bill Bradley", "count": 41 }, { "value": "Sally Jones", "count": 17 }],
    "places": [{ "value": "New York, NY, USA", "count": 209 }],
    "tags":   [{ "value": "snow", "count": 12 }]
  }
}

facets is present only when requested. In AI mode with a semantic query, facets are computed over the filter scope (tenant + advancedFilters) rather than the ranked result set, because facets are counts, not rankings. Values are exact strings as indexed (case-sensitive).

[!NOTE] Facets are available from API version 1.0.78. Assets indexed before the upgrade appear in the faces, labels, places, and tags facets after the library's one-time reindex; mimeType and camera facets work immediately.

Similar Assets

Search answers "find assets matching these words". Similar assets answers "find assets that look like this one". Every asset ingested with vectorize: true (the default) has an embedding in the vector index, and GET /assets/:id/similar returns the nearest neighbours of that embedding, most similar first.

GET /assets/5e92fe62-40a4-42fe-aa99-0a2d84a91cb7/similar?limit=12
ParameterTypeNotes
limitint, 1–50, default 12Number of matches returned
minScorenumber, 0–1Optional. Drops matches scoring below it. Omitted, the nearest limit are returned whatever they score
{
  "assetId": "5e92fe62-40a4-42fe-aa99-0a2d84a91cb7",
  "results": [
    { "id": "123e4567-e89b-12d3-a456-426614174000", "matchScore": 0.91, "originalName": "harbor-at-dusk.jpg", "mimeType": "image/jpeg" }
  ],
  "total": 1
}

The embedding is read from the index, so no model is called at request time. The asset itself and trashed assets are never returned, and the scope is the caller's tenant, as in search.

Filtering the Matches

POST /assets/:id/similar takes limit and minScore in the body, plus the advancedFilters and excludeAssetIds that POST /search accepts. Use it to compare pictures only with pictures, or to keep the matches within what your own user may see.

POST /assets/5e92fe62-40a4-42fe-aa99-0a2d84a91cb7/similar
{
  "limit": 24,
  "advancedFilters": { "type": "image", "metadata": { "asset_status": "Active" } },
  "excludeAssetIds": ["9b2e6c1a-7d43-4f0e-8a55-1c2d3e4f5a6b"]
}

The filters are applied inside the nearest-neighbour search. You receive the limit nearest assets of the filtered scope, not a filtered remainder of the unfiltered nearest.

Assets Without an Embedding

An asset that was ingested with vectorize: false, or is still processing, has nothing to compare. The response is an empty list with a reason, not an error:

{ "assetId": "5e92fe62-40a4-42fe-aa99-0a2d84a91cb7", "results": [], "total": 0, "reason": "no-embedding" }

[!NOTE] Similar assets are available from API version 1.0.94.