- Carry unmanaged stored config keys when saving, OAuth, and device auth - Add tests for credential carryover and cleared optional fields - Document carriedStoredConfig behavior in architecture docs
243 KiB
Songhive Architecture
System Overview
Songhive is a federated and self-hosted music sharing service. It uses ActivityPub for federation (via pubby) and is designed to interoperate with Mastodon-compatible clients and the fediverse at large.
┌─────────────────────────────────────────────────────────────────────────────┐
│ External Services │
│ ┌──────────────┐ ┌────────────────┐ ┌─────────────────────────────────┐ │
│ │ MusicBrainz │ │ Cover Art │ │ Federated Instances (fediverse) │ │
│ │ (metadata + │ │ Archive (cover │ │ (ActivityPub inbox/outbox) │ │
│ │ MBID lookup)│ │ art images) │ └─────────────────────────────────┘ │
│ └──────────────┘ └────────────────┘ │
└─────────────────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ Reverse Proxy (Nginx) │
└─────────────────────────────────────────────────────────────────────────────┘
│
┌─────────────────┼─────────────────┐
▼ ▼ ▼
┌─────────────────────────┐ ┌─────────────────┐ ┌─────────────────────────────┐
│ Frontend (Vue.js 3) │ │ Tornado Server │ │ Static/Media Files │
│ TypeScript + Vite │ │ (FastAPI ASGI) │ │ (Local filesystem / S3) │
└─────────────────────────┘ └─────────────────┘ └─────────────────────────────┘
│
┌───────────────────────┼──────────────────────┐
▼ ▼ ▼
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ PostgreSQL │ │ Redis │ │ Celery │
│ (Primary │ │ (Cache / │ │ (Workers) │
│ Database) │ │ Broker / │ │ │
│ │ │ Sessions) │ │ │
└─────────────┘ └─────────────┘ └─────────────┘
Backend Stack
| Component | Technology |
|---|---|
| API Framework | FastAPI (ASGI) |
| Server | Tornado (wrapping FastAPI via a2wsgi) |
| Fallback server | uvicorn (when a2wsgi is unavailable) |
| WebSockets | Tornado native WebSocket handlers |
| Data Models | Pydantic v2 (validation/serialization) |
| ORM | SQLAlchemy 2 (async, mapped columns) |
| Task Queue | Celery + Redis |
| Cache/Sessions | Redis |
| Federation | pubby (ActivityPub library) |
| Auth | JWT access tokens + opaque refresh tokens (Redis) |
| OAuth2 | authlib (OAuth2 provider for third-party apps) |
SMTP (via Python smtplib) |
|
| Metadata | mutagen (tag reading), MusicBrainz API |
| Transcoding | ffmpeg (system dependency) |
| Rate limiting | Redis sliding-window (per-IP and per-user) |
Application Structure
songhive/
├── __init__.py
├── __main__.py
├── app.py # Entry point: Tornado+FastAPI bootstrap; uvicorn fallback
├── version.py
├── cli/ # Admin commands (init-db, migrate, create-user, etc.)
│ └── admin.py
├── config/ # Configuration management
│ ├── schema.py # Pydantic BaseSettings model (all subsections)
│ └── loader.py # TOML + env vars + CLI argument loading
├── api/ # FastAPI application
│ ├── app.py # FastAPI factory: middleware, route registration, federation setup
│ ├── _common.py # Shared helpers (pagination, client IP, etc.)
│ ├── cookies.py # Set/clear HttpOnly auth cookies + double-submit CSRF cookie
│ ├── deps.py # Dependency injection (DB session, current user, config, Redis, storage)
│ ├── errors.py # RFC 7807 problem-detail exception handlers
│ ├── semantic_meta.py # OpenGraph/rel="tag" <head> tag builders + SPA shell injection
│ ├── routes/ # Route modules (one file per resource)
│ │ ├── auth.py # Login, registration, token refresh, password reset
│ │ ├── sessions.py # List and revoke active refresh-token sessions
│ │ ├── users.py # Public user profiles, directory, per-user activity feed; authenticated profile updates
│ │ ├── profile_pages.py # Browser SPA profile routes, rel="me" link injection, ActivityPub negotiation
│ │ ├── activities.py # Activity interactions (like, edit, delete, …)
│ │ ├── statuses.py # Standalone status composer endpoint (POST /statuses)
│ │ ├── artists.py
│ │ ├── albums.py
│ │ ├── tracks.py
│ │ ├── playlists.py
│ │ ├── libraries.py
│ │ ├── favorites.py
│ │ ├── tags.py # Global tag browsing and admin deletion
│ │ ├── genres.py # Global genre browsing and admin deletion
│ │ ├── history.py # Listening history
│ │ ├── stats.py # Personal listening statistics
│ │ ├── radios.py # Dynamic radio generation
│ │ ├── files.py # Generic file upload/download (StoredFile)
│ │ ├── downloads.py # Bulk-download archive requests (ZIP), status, file download, clear
│ │ ├── shares.py # Share grants (owner→specific user) + /shares/mine listing
│ │ ├── share_urls.py # Share URL tokens (revocable short links)
│ │ ├── share.py # Public token resolver (redirects + sets cookie)
│ │ ├── reports.py # Content moderation reports + admin review
│ │ ├── feeds.py # RSS/Atom feeds under /feeds (users, entities, tags, genres)
│ │ ├── federation.py # ActivityPub object endpoints (tracks, objects, federated music entities) + WebFinger
│ │ ├── remote.py # Explicit remote lookup/dereference + cached remote objects
│ │ ├── admin.py # Admin endpoints (settings, stats, user management)
│ │ ├── external_libraries.py # User external library CRUD, sync, tracks
│ │ └── admin_external_libraries.py # Admin external library management
│ └── middleware/
│ ├── auth.py # JWT decode middleware + access-token helpers
│ ├── csrf.py # Double-submit CSRF check for cookie-authenticated unsafe requests
│ ├── head.py # HEAD answered as bodyless GET (crawler probes, e.g. og:image)
│ ├── media_cors.py # Wildcard CORS on read-only media endpoints (federation embeds)
│ ├── rate_limit.py # Redis sliding-window rate limiting (IP / user)
│ └── proxy.py # X-Forwarded-Proto scheme handling behind reverse proxies
├── migrations/ # Alembic database migrations
│ ├── env.py # Migration environment (imports all models)
│ ├── script.py.mako # Template for generated revisions
│ └── versions/ # Revision scripts
├── models/ # SQLAlchemy mapped models + shared enums
│ ├── base.py # DeclarativeBase, UUID PK, timestamps, async session factory
│ ├── _enums.py # Visibility enum (private / mentioned / local / followers / public)
│ ├── activity.py # Activity, ActivityMention, ActivityTarget (federation interaction layer)
│ ├── preview_card.py # PreviewCard — per-URL OpenGraph/<title>/domain card cache
│ ├── remote_object.py # RemoteObject — cached remote AP objects + normalized fields
│ ├── user.py # User (roles: user / moderator / admin; federation fields)
│ ├── user_link.py # Profile links (validated URL list)
│ ├── invite.py # Invite codes (max_uses, expiry)
│ ├── artist.py
│ ├── album.py
│ ├── track.py
│ ├── upload.py # Raw uploaded file reference
│ ├── stored_file.py # Content-addressable file (SHA-256, visibility, owner)
│ ├── transcoded_file.py # Transcode cache: (track_id, format, bitrate) → StoredFile
│ ├── download.py # DownloadArchive — async bulk-download ZIP request (items snapshot, item_errors)
│ ├── library.py
│ ├── library_track.py # Library ↔ Track join table
│ ├── playlist.py
│ ├── favorite.py
│ ├── follow.py # Outbound follow relationships (local users → local/remote actors)
│ ├── genre.py # Genre and GenreTrack/GenreAlbum associations
│ ├── history.py # Listening history entries
│ ├── radio.py
│ ├── share_grant.py # Per-user access grant for a specific item
│ ├── share_token.py # Revocable short-link token (stores SHA-256 hash only)
│ ├── report.py # Content moderation report
│ ├── oauth_client.py # Registered OAuth2 clients
│ ├── audit_log.py # Admin/security audit trail
│ ├── external_library.py # External library adapter instance
│ ├── external_sync_run.py # External library sync history
│ ├── external_track.py # Track discovered through an external provider
│ ├── output_stream.py # Server-side audio output (encrypted provider config)
│ ├── playback_session.py # Persisted playback control plane for server outputs
│ └── setting.py # Runtime-editable instance settings (key/JSON-value)
├── services/ # Business logic layer
│ ├── acl.py # Three-level visibility + share-grant + share-token ACL
│ ├── activities.py # Activity domain service (entity resolution, creation, interactions)
│ ├── auth.py # User lookup, password hashing, session helpers
│ ├── audit.py # Audit log helpers
│ ├── deletion.py # Cascade deletion + activity retraction fan-out
│ ├── downloads.py # Archive item resolution (ACL), ZIP materialization, name sanitization
│ ├── email.py # SMTP email (verification, password reset)
│ ├── federation.py # Actor provisioning, domain allow/block, inbox dispatch
│ ├── feeds.py # RSS 2.0/Atom feed documents + per-entity feed queries
│ ├── follows.py # Outbound follow/unfollow, decision folding, follow listings
│ ├── genres.py # Genre validation, association and listing
│ ├── import_.py # Import pipeline orchestration
│ ├── listening_stats.py # Per-user listening stats aggregation (top-N, histograms, clock)
│ ├── mentions.py # @handle extraction, local/WebFinger resolution, safe HTML rendering
│ ├── remote_content.py # Remote lookup parser, actor/object dereference + cache, cached search
│ ├── metadata.py # Tag extraction coordination
│ ├── music.py # Music library helpers
│ ├── musicbrainz.py # MusicBrainz + Cover Art Archive enrichment (async httpx)
│ ├── outputs.py # Output CRUD, config encryption, mount resolution
│ ├── playback.py # PlaybackSession state machine + worker commands
│ ├── redis.py # Redis client lifecycle
│ ├── reports.py # Content report CRUD
│ ├── secrets.py # Config encryption, decryption, and secret redaction
│ ├── settings.py # Instance settings with Redis cache + config overlay
│ ├── sharing.py # Share-grant and share-token CRUD
│ ├── stats.py # Admin dashboard statistics
│ ├── storage.py # StorageService facade (delegates to storage backend)
│ └── streaming.py # Track file resolution, transcode cache, history recording
├── external/ # External library adapter interface and registry
│ ├── base.py # BaseExternalLibraryAdapter protocol
│ ├── _fake.py # In-memory fake adapter for tests
│ ├── _local.py # Local filesystem adapter
│ ├── watchdog.py # Filesystem watchdog for local libraries
│ ├── registry.py # Provider-type registry
│ └── types.py # Adapter dataclasses (ItemRef, TrackMetadata, etc.)
├── federation/ # ActivityPub per-user federation
│ ├── _common.py # URL builders
│ ├── fetch.py # SSRF-guarded remote fetch (DNS/IP checks, redirect revalidation) + guarded binary download
│ ├── actors.py # Actor document generation, federation storage helpers
│ ├── activities.py # Activity creation (Create, Update, Delete, etc.)
│ ├── doc_cache.py # Shared pubby.cache document cache: tuple-key schema + after-commit invalidation helpers
│ ├── incoming.py # Materialize inbound remote replies into Activity rows
│ ├── notifications.py # Inbox recipient resolution + notification hooks
│ ├── serializers.py # Track → ActivityPub Audio object mapping
│ └── storage.py # pubby storage adapter (SQLAlchemy-backed)
├── users/ # User management
│ ├── manager.py # User CRUD, password management
│ ├── invites.py # Invite-code creation, validation, consumption
│ ├── oauth.py # OAuth2 provider setup (authlib)
│ └── tokens.py # JWT access + opaque refresh token issuance/rotation/revocation, session listing
├── music/ # Music domain logic
│ ├── importer.py # File importer: save → extract tags → link track
│ └── metadata.py # Tag reading (mutagen), field normalization
├── streaming/ # Audio streaming
│ ├── handler.py # Tornado streaming handler (range requests, send-file)
│ ├── mount.py # Tornado listener handler for native HTTP mounts
│ └── transcoder.py # ffmpeg wrapper: MP3, OGG, FLAC, AAC, Opus
├── streams/ # Server-side audio outputs
│ ├── base.py # AudioOutput provider base class (fields, redaction)
│ ├── driver.py # OutputDriver runtime interface (start/pause/seek/events)
│ ├── registry.py # provider_type → AudioOutput registry
│ ├── icecast.py # Icecast provider: ffmpeg → icecast:// pipeline
│ ├── http.py # Native HTTP provider: ffmpeg → Redis Stream fan-out
│ ├── snapcast.py # Snapcast provider: ffmpeg PCM → snapserver FIFO/TCP
│ ├── fake.py # In-memory provider for tests
│ ├── types.py # AudioSource and related dataclasses
│ └── worker.py # `songhive stream-worker` session driver loop
├── storage/ # Media storage backends
│ ├── base.py # Abstract StorageBackend interface + FileSizeLimitExceededError
│ ├── exc.py # Storage-layer exceptions
│ ├── local.py # Local filesystem backend
│ └── s3.py # S3-compatible object storage backend
├── tasks/ # Celery task definitions
│ ├── celery.py # Celery app factory (crontab parser, config loading)
│ ├── import_.py # Import pipeline tasks
│ ├── federation.py # Activity delivery + inbox processing tasks
│ ├── transcoding.py # Pre-transcoding tasks
│ ├── email.py # Email delivery tasks
│ ├── musicbrainz.py # MusicBrainz metadata + Cover Art Archive enrichment
│ ├── images.py # Artist image + album cover enrichment
│ ├── external_libraries.py # External library sync task
│ ├── preview_cards.py # Link-preview fetch + per-URL cache task
│ ├── downloads.py # Async ZIP archive builds (Redis concurrency cap) + retention cleanup (scheduled)
│ └── storage.py # Orphaned-file cleanup (scheduled via crontab)
├── ws/ # WebSocket support
│ └── events.py # Tornado WebSocket handler (JWT auth, CORS origin check)
└── cli/ # CLI commands
└── admin.py # Admin commands (create-user, provision-federation-keys, etc.)
Server Architecture: Tornado + FastAPI
The server uses Tornado as the top-level HTTP server. The FastAPI ASGI app is
bridged via a2wsgi (ASGI→WSGI) and wrapped in Tornado's WSGIContainer.
When a2wsgi is not installed the server falls back to uvicorn (pure ASGI,
loses native Tornado handlers).
Why Tornado as the outer server:
- Native WebSockets — the real-time events endpoint is a proper Tornado
WebSocketHandler, avoiding ASGI WebSocket complexity. - Streaming handler — audio files are streamed via a dedicated Tornado handler with native range-request support.
- Signal handling / graceful shutdown — Tornado's
IOLoopcontrols the process lifecycle.
# Simplified bootstrap (songhive/app.py)
from a2wsgi import ASGIMiddleware
from tornado.web import Application, FallbackHandler
from tornado.wsgi import WSGIContainer
from tornado.httpserver import HTTPServer
from songhive.api.app import create_app
from songhive.ws.events import EventWebSocket
from songhive.streaming.handler import StreamHandler
from songhive.streaming.mount import StreamMountHandler
fastapi_app = create_app(config)
wsgi_app = ASGIMiddleware(fastapi_app)
container = WSGIContainer(wsgi_app, executor=ThreadPoolExecutor(...))
tornado_app = Application([
(r"/ws/events", EventWebSocket),
(r"/ws/", EventWebSocket),
(r"/api/v1/stream/(?P<track_id>[^/]+)", StreamHandler),
(r"/streams/(?P<mount>[^/]+)", StreamMountHandler),
(r".*", FallbackHandler, {"fallback": container}),
])
server = HTTPServer(tornado_app)
server.listen(config.server.port)
Note:
WSGIContainermust be given an explicitThreadPoolExecutor. Without one (Tornado < 7 default) the WSGI app runs on the Tornado event loop thread, serializing every request — and any server-side remote fetch that triggers a synchronous call-back to this instance (e.g. a remote ActivityPub server resolving ourkeyIdfor HTTP signature verification) deadlocks, since the inbound request can never be served while the loop is busy inside the outbound lookup.
Configuration
Configuration is loaded with the following priority (highest first):
- Environment variables (prefixed
SONGHIVE_, nested with__) - CLI arguments
config.toml(searched at the path given by--configorSONGHIVE_CONFIG, then./config.toml, then$XDG_CONFIG_HOME/songhive/config.tomlor~/.config/songhive/config.toml, and finally/etc/songhive/config.toml)- Field defaults
The root schema is SonghiveConfig (a Pydantic BaseSettings), composed of
these subsections:
| Section | Key settings |
|---|---|
server |
host, port, num_workers, debug, cors_origins |
database |
url (asyncpg), pool_size, max_overflow, pool_timeout, pool_recycle, pool_pre_ping |
redis |
url |
celery |
broker_url, result_backend, cleanup_orphaned_files_schedule |
storage |
backend (local/s3), local_path, s3_*, cdn_prefix, max_upload_size, upload_quota |
federation |
enabled, instance_domain, instance_name, contact_name/contact_email/contact_url, private_key_path, allow/block lists, remote_search_access, fetch_timeout_seconds, remote_activity_retention_days, remote_activity_prune_schedule, document_cache_ttl_seconds |
auth |
registration_mode, secret_key, token TTLs, rate_limit, trusted_proxy_hops, cookie_secure, cookie_samesite, cookie_domain |
email |
smtp_host, smtp_port, smtp_user, from_address, tls settings |
musicbrainz |
enabled, user_agent, cover_art, artist_image settings |
notifications |
retention_days, purge_hour, digest_hour |
imports |
scan_roots, bulk_import_sync_threshold |
downloads |
enabled, max_items, max_concurrent_archives, max_active_per_user, retention_hours, stale_run_hours, fetch_attempts, fetch_backoff_seconds, fetch_timeout_seconds, max_item_bytes |
streaming |
max_bitrate, max_bitrate_by_role, default_bitrate, chunk_size, transcode_cache_enabled |
streams |
enabled, allow_user_created_outputs, allowed/denied_user_providers, allowed_output_hosts, icecast_ffmpeg_path, worker timings, idle timeout, http_stream_* |
Runtime-editable overrides (instance settings stored in the settings DB
table, cached in Redis) are applied over the file-based config at startup and
can be changed via the admin API without a restart.
Data Model (Core Entities)
┌─────────────┐ ┌─────────────┐ ┌─────────────┐ ┌──────────────┐
│ Artist │────▶│ Album │────▶│ Track │────▶│ StoredFile │
└─────────────┘ └─────────────┘ └─────────────┘ │ (audio_file) │
│ └──────────────┘
│
▼
┌─────────────┐ ┌──────────────┐
│ Upload │────▶│ StoredFile │
│ (raw upload)│ │ (upload file)│
└─────────────┘ └──────────────┘
│
▼
┌──────────────────┐
│ TranscodedFile │
│ (track, format, │
│ bitrate) │
└──────────────────┘
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ User │────▶│ Library │────▶│ LibraryTrack│
│ (roles: │ └─────────────┘ └─────────────┘
│ user / │
│ moderator /│ ┌─────────────┐ ┌─────────────┐
│ admin) │────▶│ Playlist │────▶│ Track │
│ │ └─────────────┘ └─────────────┘
│ │
│ │ ┌─────────────┐ ┌─────────────┐
│ │────▶│ ShareGrant │────▶│ (any item) │
│ │ └─────────────┘ └─────────────┘
│ │
│ │ ┌─────────────┐ ┌─────────────┐
│ │────▶│ ShareToken │────▶│ (any item) │
│ │ └─────────────┘ └─────────────┘
│ │
│ │ ┌─────────────┐ ┌─────────────┐
│ │────▶│ Invite │ │ UserLink │
│ │ └─────────────┘ └─────────────┘
└─────────────┘
Visibility levels (Visibility enum) are ordered from most to least
restrictive: private < mentioned < local < followers < public
(Visibility.rank, Visibility.can_contain). Entity visibility currently
uses private, local, and public (applies to tracks, albums, artists,
libraries, stored files):
private— visible only to the owner (and users with aShareGrant)local— visible to authenticated users on the same instancepublic— visible to everyone including federated instances
The mentioned and followers levels exist for the activities layer
(Activity.visibility) and are not yet meaningful for entity ACLs.
Activities
Activity records federation-relevant events (create, announce, like,
reply, quote, mention, update, delete, webmention) attached to an
entity through (entity_type, entity_id) — where entity_type is one of
track, album, artist, playlist, library, user (the user entity
hosts standalone statuses posted through POST /api/v1/statuses), remote
(remote objects dereferenced on demand — see below). Each row
tracks its origin via source_type (local or a remote source),
source_actor, and source_id (unique per source), with local_object_id
as an optional canonical local identifier. Activities support threading
through in_reply_to_activity_id, arbitrary JSON payloads, source text vs
rendered content (content_source / content / content_type — plain text
or Markdown), an optional BCP-47 language tag mirrored into the object's
contentMap, and soft deletion (deleted_at, retracted). An activity's
visibility must not exceed its parent entity's visibility
(Visibility.can_contain).
ActivityTag rows link Activity and Tag, populated whenever an activity
contains hashtags (#tag) in its content_source/content. They enable
GET /api/v1/tags/{tag}/activities, which returns tag-matching activities
subject to the same visibility and entity ACL rules as other activity feeds.
The same rows are folded into the tag listing queries in
services/tags.py, so GET /api/v1/tags/ and GET /api/v1/tags/{tag}
also surface hashtags that appear only on activities (as activity items),
counted by published_at and filtered by activity visibility plus the
containing entity's ACL.
Remote ActivityPub content is processed by Pubby and is mostly not
materialized as local Activity rows — the exception is inbound Create
replies, which federation/incoming.py stores as source_type="remote"
rows (see below) so their hashtags and mentions are associated too. Other
remote content stays interaction-only, so hashtags from remote posts cannot
be associated on the Songhive side unless Pubby stores or forwards them.
ActivityMention rows capture @handle mentions embedded in content, with
optional actor_url / user_id resolution and a notified_at marker.
services/mentions.py implements the mention pipeline: MENTION_REGEX
extracts @user and @user@domain handles, resolve_mentions resolves bare
handles against the local users table (case-insensitive, active users only)
and remote handles through pubby.resolve_actor_url (run in a thread, with
Pubby's https://{domain}/@{username} fallback on lookup failure). Remote
resolution is gated on federation.enabled, requires a dotted domain, and
drops handles on blocked or non-allowed instances via
services/federation.is_domain_blocked; @user@domain handles naming the
local instance resolve locally instead. render_mentions builds safe HTML —
resolved handles become anchors via pubby.render_link_anchor, surrounding
text is escaped and linkified by pubby.render_post_html (which also
converts newlines to <br>, since remote servers render content/summary
as HTML and would collapse literal newlines) — and
process_mentions is the single entry point returning resolved mentions,
rendered HTML, tags, and ActivityPub Mention/Hashtag tags. It accepts a
content_type: text/plain uses the escaped plain-text renderer described
above, while text/markdown (render_mentions_markdown, the default for new
statuses) runs the source through mistune — with raw HTML escaped, bare
URLs linkified, javascript:-style link targets neutralized — before the
same mention/hashtag linkification applies; handles inside code spans are
left untouched.
ActivityTarget rows track per-inbox outbound delivery state (pending,
sent, failed, skipped) with attempts / last_error /
last_attempt_at bookkeeping. Tracks already published to the fediverse
(federation_object_id set) are backfilled as create activities by
migration 4adb5fbea9d6.
Preview cards attach link-preview metadata to activities: when a post is
created locally (create_status, reply_to_activity, quote_activity,
record_track_publication) or materialized from a remote Create,
schedule_preview_card_fetch enqueues tasks/preview_cards.py's
fetch_preview_card, which runs services/preview_cards.py's
process_activity_preview_card. The service picks the first pure URL —
Mastodon-style — from the raw content_source of local posts (rendered
mention/hashtag anchors never appear there) or from the anchors of the
remote HTML content, skipping mention/hashtag/u-url classes and
rel="tag" links, then fetches OpenGraph metadata with <title> and
finally the URL's domain as fallbacks. Fetched cards live in the
preview_cards table keyed by normalized URL and are shared across
activities; a cached card is reused while fresher than
PREVIEW_CARD_MAX_AGE (24h) and re-fetched at post time otherwise, so
views never trigger network fetches. Activities whose object carries
attachments get no card, and edits re-run the pipeline (force=True) so a
stale link is refreshed or cleared. Fetches are guarded against SSRF —
only http(s) URLs resolving to globally routable addresses are
requested, redirects are re-validated per hop, and bodies are capped at
MAX_DOCUMENT_BYTES (1 MiB). The pipeline is gated by the instance-level
preview_cards_enabled runtime setting (admin UI, default on) and by the
per-user preview_cards_enabled preference (profile form, default on);
either opting out skips the enqueue, and the task re-checks both gates so
rows created before a change are honoured. ActivityResponse.preview_card
serializes the linked card and ActivityCard.vue renders it below the
post content.
services/activities.py is the domain entry point: resolve_entity maps an
(entity_type, entity_id) pair to its model row, and
create_local_activity validates the entity (404), enforces
Visibility.can_contain against the entity (422 — entities without a
visibility column, currently Artist, act as public containers), checks
acl.can_manage (403), then persists the activity plus any pre-resolved
ActivityMention rows. Local source identity reuses the author's
actor_url, falling back to a urn:songhive:user:{username} URN when the
user has no provisioned federation identity; source_id embeds the same
UUID as local_object_id ({actor}/objects/{uuid}) so section-9 object
routes can resolve it.
can_view_activity decides who may see an activity: the containing entity
must pass acl.can_access, and the activity's own visibility applies on
top (public follows the entity check; local/followers require an
authenticated user; mentioned requires the owner or a mentioned user;
private is owner-only; retracted activities are never viewable). Admins
get no bypass — mentioned and private activities stay confined to
their audience for every viewer. The same rules are enforced in SQL by
_activity_visibility_filter in services/acl.py (feed/tag/reply
listings) and in
reply_count computation, so restricted replies neither render nor leak
through counters for viewers outside their audience. Interactions are layered on top of create_local_activity:
like_activity records an idempotent like (400 on a duplicate, 404 on a
retracted target) that inherits the target's visibility and stores a
federation/activities.create_like_activity Like payload whose to/cc
come from activity_audience. boost_activity mirrors it as an
announce activity carrying an Announce payload
(federation/activities.create_announce_activity, a thin wrapper over
pubby's build_announce_activity) whose object is the target's
source_id; duplicates are rejected the same way. reply_to_activity
stores a reply activity whose payload is a Create(Note) with
inReplyTo set to the target's source_id — the reply reuses the status
composition pipeline (process_mentions, media/track attachment
resolution, contentMap language tagging) and its requested visibility
may never exceed the target's (Visibility.can_contain) or the entity's.
The replied-to author is always addressed: a local owner gets an
ActivityMention row plus a reply notification (_notify_reply), a
remote author gets a Mention tag so resolve_audience delivers the
reply to their inbox. Interactions skip the can_manage gate
(require_manage=False) — view access on the target is enough — while
content-producing activity types still require manage rights. The
POST /api/v1/activities/{id}/like, /{id}/boost, /{id}/reply and
/{id}/quote
endpoints perform the 404/403/400 (or 422 for invalid replies) checks,
provision the author's actor keys (ensure_user_actor), commit, then fan
out: likes and boosts share _fan_out_reaction_activity (exposed as
fan_out_like_activity/fan_out_boost_activity), which — for remote
targets — resolves the reacted author's inbox via
services/federation.resolve_actor_inbox (the federation_actor_cache
table first, then a signed actor-document fetch, preferring sharedInbox)
and hands it to fan_out_activity alongside the reaction's own audience;
replies fan out inline through fan_out_activity, best-effort. Quotes
(quote_activity) work like replies — a Create(Note) row linked to the
quoted activity through in_reply_to_activity_id — but the note carries
the FEP-0449 quote field plus the quoteUri/quoteUrl forms Fedibird,
Akkoma and Mastodon read and Misskey's _misskey_quote (never
inReplyTo: a quote is not a reply), the quoted author is always
addressed via a Mention tag — and activity_audience puts mentioned
actors in to for public and followers posts, matching how Mastodon and
Akkoma address their own — and FEP-044f authorization
is negotiated: quoting a local user's post self-issues the
QuoteAuthorization their auto-approval would grant (stored so it
dereferences, stamped as quoteAuthorization on the note), while quoting
a remote post delivers a QuoteRequest — quoting Note as instrument,
quoted object as object — to the remote author's inbox
(_deliver_quote_request, sent before the Create fan-out so the
authorization flow starts first). The remote Accept answering it stamps the
issued authorization id onto the stored quote (see inbound
materialization below). DELETE /api/v1/activities/{id}/like and
/{id}/boost retract the
caller's reaction: unreact_activity soft-deletes the like/announce row
(so the target can be reacted to again), removes the notification it
produced, and returns the retracted activity; fan_out_unreaction_activity
then delivers an Undo wrapping the originally federated Like/Announce
(federation/activities.create_undo_activity, over pubby's
build_undo_activity) to exactly the inboxes recorded as sent for the
reaction. Unreacting a target without a live reaction returns 404.
Every ActivityResponse carries interaction state resolved by
resolve_interaction_summaries: like_count and boost_count combine
direct local Activity rows (interactions stored with
in_reply_to_activity_id pointing at the target) with confirmed remote
interactions recorded in Pubby's federation_interactions storage against
the activity's source_id (_remote_interactions, thread-offloaded and
best-effort). reply_count instead covers the whole sub-thread: every
non-deleted reply descendant reachable through the in_reply_to chain —
computed with the _activity_descendants_cte recursive CTE — plus remote
replies targeting any node in it, including remote replies to remote
replies reached breadth-first through object_id chains
(_remote_thread_interactions). quote_count is flat — quotes attach to
the quoted post itself, never to each other — and counts direct
quote-type children (local quotes and materialized remote ones) plus
confirmed remote QUOTE interactions, deduplicated on object_id.
Remote replies already materialized into
Activity rows are skipped so a reply backed by both a row and a stored
interaction counts once. Descendants the requester may not see
(mentioned replies naming someone else, private rows owned by another
user) are not counted, though traversal still crosses them to reach their
visible children. liked/boosted report the requester's
own live like/announce rows, and can_interact is false for activity types
that cannot themselves be reacted to (like, announce, delete).
ActivityResponse also carries object_url/object_type, resolved from
the payload's object: Create-style activities expose the embedded
document's id/type, while Like/Announce payloads reference the
reacted object as a bare id — which is exactly the target a reaction card
should link to — and expose no object_type. PATCH rejects like and
announce activities (422): a reaction's object is a bare reference and
its visibility is inherited, so there is nothing editable. GET /api/v1/activities/{id}/likes and /{id}/boosts return the
known interactors — local users resolved to profiles, remote actors from
Pubby's interaction records — via list_activity_interactors, and
GET /api/v1/activities/{id}/replies returns the known replies via
list_activity_replies: visibility-filtered local reply activities —
every descendant in the thread, not just direct children — serialized as
full ActivityResponses plus remote replies rebuilt from the raw_object
metadata the inbox processor stored, each carrying in_reply_to (the
replied-to object id) so clients can regroup the flat list into threads.
Replies already materialized into Activity rows are filtered out of the
remote list so they render only as full cards. GET /api/v1/activities/{id}/quotes lists the known quotes via
list_activity_quotes: visibility-filtered local quote activities —
locally authored ones and remote quotes materialized into Activity
rows — serialized as full ActivityResponses, plus remote quotes rebuilt
from Pubby InteractionType.QUOTE records keyed by the quoted object's
source_id, each carrying quoted (the quoted object id) instead of the
reply's in_reply_to. Materialized remote quotes are filtered out of the
remote list so they render only as cards.
The frontend renders them Mastodon-style: every descendant of a direct
reply is unfolded flat into that reply's thread, each thread marked by its
own vertical line. All four listings gate on can_view_activity and allow
anonymous reads of public targets.
GET /api/v1/activities/{id} returns a single activity under the same
rules — it backs the notification UI's embedded activity cards and the
SPA's /activities/:id permalink page (views/ActivityView.vue), which
also lists the known quotes (GET /{id}/quotes) under the card — local
ones as ActivityCards, federated-only ones as ActivityRemoteReply
rows — failure-isolated so a listing error cannot sink the page.
ActivityCard.vue carries a separate Quote action next to Reply: a
StatusComposer posts the quote through POST /api/v1/activities/{id}/quote, the quote_count button toggles a flat
chronological list of local and remote quotes, and a quote card itself
embeds the quoted post via ActivityObjectEmbed.
GET /api/v1/activities/lookup?url=... resolves an ActivityPub object URL
to the same ActivityResponse: it matches Activity.source_id directly
(covering materialized remote replies whose object ids live on remote
paths) and, for {actor}/objects/{uuid} shapes, local_object_id. Browser
requests to a local /users/{name}/objects/{id} that map to a stored
activity are redirected by federation.py to the /activities/{id}
permalink, while ActivityPub Accept headers still get the object JSON.
Front proxies that serve the SPA shell for browser traffic bypass that
redirect, so the SPA also mounts a matching
/users/{username}/objects/{objectId} route
(views/ObjectPermalinkView.vue) that resolves the permalink client-side
through /api/v1/remote/lookup (which maps local URLs to their SPA route
without any remote fetch, and is therefore exempt from the
remote_search_access policy for local targets) with an
/api/v1/activities/lookup fallback.
Notification payloads denormalize object_activity_id/object_type/
object_page_url for the note itself and target_object_* for the
replied-to/quoted activity, so clients can render real activity cards and
link replies to their parent instead of a remote object id that does not
dereference to a page.
resolve_audience maps an activity's visibility to the set of remote
inbox URLs it should reach: public and followers activities go to
the author's follower inboxes (services/federation.get_follower_inboxes
reads Pubby's federation_followers storage via pubby.collect_inboxes,
preferring shared_inbox and deduplicating) plus every remote mentioned
actor, while mentioned activities reach only the mentioned actors.
public activities additionally reach followers of objects in the
reply chain — remote actors can Follow a local object rather than an
actor (Friendica sends Follow on a thread's root item for conversation
subscriptions; see Object follows below), and
services/federation.get_object_follower_inboxes collects the inboxes
subscribed to the activity's own object id or any of its
in_reply_to_activity_id ancestors.
Mentions that resolved to a local user (user_id set) or an actor URL on
the local instance domain have no remote inbox and are excluded, as are
non-HTTP(S) actor URLs; private and local never federate. Mentioned
inboxes are resolved concurrently through resolve_actor_inbox, signed
with the owner's key.
fan_out_activity is the delivery step: it unions the resolved audience
with caller-supplied extra_inboxes, drops inboxes already booked, and
creates an ActivityTarget row per remaining inbox before enqueueing
tasks.federation.deliver_activity (which applies the federation gate, the
blocked-domain gate, and exponential-backoff retries). Targets are marked
sent once the task accepts the job, failed with last_error when
enqueueing raises, and skipped — without an attempt — for blocked or
non-allowed inbox domains; pending remains the default for rows staged
without a dispatch attempt. Fan-out no-ops for remote or retracted
activities, missing payloads, non-federating visibilities, disabled
federation, and owners without a signing key.
VisibilityRules centralizes the containment policy: can_contain and
enforce_activity_visibility (used by create_local_activity) validate an
activity's visibility against its entity's, while
cascade_visibility_update changes a local activity's visibility and fans
the change out to the inboxes recorded as sent in activity_targets.
When the new visibility still federates (Visibility.federates:
mentioned, followers, public) those inboxes receive an Update
carrying the new to/cc audience built by
federation.activities.create_visibility_update_activity on top of
activity_audience; when it does not (private, local) a
Delete(Tombstone) retracts the object instead. Remote activities and
soft-deleted local activities never fan out, and deliveries are signed with
the activity owner's key through deletion.enqueue_activity_delivery.
PATCH /api/v1/activities/{id} handles author edits behind an
acl.can_manage check on the containing entity (owner or admin — the same
gate create_local_activity applies to content-producing types). Only
fields present in the request body change: a content edit goes through
services.activities.update_activity, which re-runs the
process_mentions pipeline on the new content_source (honoring the
stored or newly supplied content_type): content is re-rendered as
mention-aware safe HTML, the activity_mentions rows are replaced with
the newly resolved set, and — when the stored payload embeds a dict
object — the object's content and tag are rebuilt
(pubby.set_object_content merges tags while preserving pre-existing
tags, then the pipeline's Mention tags and HTML are layered on) and it
is stamped with updated. language replaces the activity's BCP-47 tag
(null clears it) and rewrites or drops the object's contentMap.
media_ids/track_ids rebuild the object's attachment list: docs
stamped with songhive:fileId/songhive:trackId (emitted by
stored_file_to_attachment/track_to_attachment) are user-managed and
replaced wholesale, while unmarked entity-owned docs — e.g. a shared
track's own Audio attachment — are preserved; the same ownership,
access, visibility-escalation and share-grant rules as status creation
apply, evaluated against the post-edit visibility. A visibility edit
goes through cascade_visibility_update so already-delivered inboxes
receive an Update or a Delete(Tombstone). Content edits then federate
through services.activities.fan_out_activity_update — invoked after both
edits so the final visibility applies — which wraps the rebuilt object in
an Update (federation.activities.create_object_update_activity, with
to/cc rewritten to the current audience) and delivers it via
deletion.enqueue_activity_delivery to every inbox recorded as sent.
fan_out_activity then re-delivers the stored payload to inboxes first
reached by the edit (e.g. a newly mentioned actor), so those recipients
get the Create carrying the updated object rather than an Update for
an object they have never seen. The fan-out no-ops for remote or
retracted activities, payloads without an embedded dict object,
non-federating visibilities, disabled federation, and owners without a
signing key.
GET /api/v1/{entity_type}/{entity_id}/activities (a second router in
api/routes/activities.py, mounted at api_prefix) is the public read
endpoint. It validates entity_type against ACTIVITY_ENTITY_TYPES
(400), resolves the entity (404), and gates on acl.can_access (403) —
anonymous requesters may read publicly accessible entities, matching the
other read endpoints; share tokens are not honored, consistent with the
like/edit routes. services.activities.list_activities then applies the
same per-activity visibility rules as can_view_activity in SQL
(_activity_visibility_filter) so pagination cannot leak or under-fill
pages, supports activity_type/source_type filters, and
keyset-paginates on (published_at, id) newest-first with an opaque
base64url cursor (limit 1–100, default 20; malformed cursors return
400). Responses are serialized through
ActivityResponse/ActivityListResponse: the raw payload,
retracted, and deleted_at stay internal while resolved mentions are
included.
Deletion is handled by services/deletion.py's cascade_delete_entity,
which is invoked from every entity delete path (track, album, artist,
playlist, library). Local activities are soft-deleted (deleted_at set)
and a Delete(Tombstone) is enqueued through deliver_activity for every
inbox recorded as sent in activity_targets, signed with the activity
owner's private key; remote activities are hard-deleted without fan-out.
Single-activity retraction follows the same rules through
services/activities.retract_activity, exposed as DELETE /api/v1/activities/{id} behind acl.can_manage on the containing entity —
except that retracting a like/announce enqueues an Undo wrapping the
originally federated reaction payload (matching the unreact_activity
path) instead of a Delete(Tombstone), since Delete is not the
ActivityPub way to retract a reaction, and remote activities are
soft-deleted rather than removed: the remote instance owns retraction, and
the surviving row keeps the stored Pubby interaction deduplicated so the
reply cannot resurface.
Track fediverse publications are recorded as activities too: every publish
path — the manual POST /api/v1/tracks/{id}/publish, uploads and imports
of public tracks that opt in with publish=true (including the
process_upload Celery task and bulk library uploads), PATCH /api/v1/tracks/{id} transitions to public that opt in with publish=true
(surfaced as a "Publish on the Fediverse" checkbox on the track metadata
edit forms; without it the transition only makes the track's URL publicly
accessible), and album visibility propagation to public — calls
services/activities.record_track_publication. Uploads are local-only by
default: without the flag the track gets no federation_object_id and no
publication activity. record_track_publication builds a Create
envelope around one of two object shapes, selected by object_type:
audio (the canonical publication — the Audio object built on the
track's freshly minted federation_object_id) or note (a share whose
content renders as the post body on every remote server and whose
Audio-typed attachment links back to the track's published object; a
fresh object id is minted per share). It stores the payload on a create
activity whose source_id matches the published object URL
({actor_url}/objects/{object_id}), keeps the
one-off status post text in content_source, persists its resolved
mentions as activity_mentions rows, records the caller-selected
visibility (default public), and delivers through
fan_out_activity so each reached inbox is booked in activity_targets
and later retraction reaches exactly those inboxes. When a public track
goes private, retract_track_publications soft-deletes all the live
publication rows — both the canonical Audio and any Note shares —
delivering a Delete(Tombstone) per publication source_id through
retract_activity, alongside the track-level
unpublish_track_activity Delete(Tombstone). Metadata edits on a
published track — PATCH /api/v1/tracks/{id} touching title,
artist_name, description, or genre — re-sync the stored object
through services/activities.sync_track_publications: each live local
create activity's payload.object is rebuilt from
track_to_audio_object or track_to_note_object depending on the stored
object type (keeping the object id stable — and the published stamp
for Note shares — while stamping
updated), a one-off status kept in content_source is re-applied as
the post body so the track description does not clobber it, and the
result fans out via fan_out_activity_update so delivered inboxes get an
Update carrying the refreshed object.
get_activity_unpublish_info returns the ActivityUnpublishInfo
(activity id, source_id, actor_url, sent inboxes) used for that
delivery, and federation/activities.py builds the Delete(Tombstone)
payload via create_tombstone_delete_activity — a thin adapter around
pubby.build_delete_activity (0.3.2) that keeps a plain string @context.
Standalone statuses — posts not attached to any media entity — are created
through POST /api/v1/statuses/, which calls
services/activities.create_status after provisioning the author's actor
(ensure_user_actor). The service records a create activity on the
author's user entity, so statuses show up in the author's profile posts
feed (GET /api/v1/users/{username}/activities?mode=posts) and can be
edited/retracted through the regular activity endpoints. Posts mode
returns the user's local create and quote activities plus —
Mastodon-style — their announce boosts unless include_boosts=false,
and their reply activities only with include_replies=true (replies
are hidden by default); mode=all returns every authored activity and
ignores the include flags. The request
carries the raw status source text (rendered through process_mentions
as text/markdown — the default — or text/plain), a visibility, an
optional BCP-47 language (validated, stored on the activity, mirrored
into the Note's contentMap), and two attachment sets: media_ids
reference previously uploaded StoredFiles (serialized by
federation/serializers.stored_file_to_attachment; the service escalates
each file's visibility to what the status audience requires — never
downgrades — and grants file shares to mentioned local users on
mentioned-visibility posts) and track_ids reference hosted tracks the
author may access (serialized by track_to_attachment as Audio
attachments embedding the stream URL, or Document links for tracks
without audio). Audio attachments — on statuses and on the Audio/Note
track objects alike — additionally carry the namespaced
songhive:trackId, songhive:trackTitle, songhive:artistName,
songhive:albumName and songhive:trackUrl keys plus a standard image
entry resolving cover art (track image, then album cover file, then the
album's remote cover_url, then the artist's image file or remote
image_url), so music-aware consumers can render a rich
player without parsing the flat name label; remote servers ignore
unknown keys. Each category is capped at four attachments, and a status
may be attachments-only. The resulting Create(Note) is built by
federation/activities.create_status_activity and fanned out through
fan_out_activity exactly like entity activities: it federates to the
visibility audience when federation is enabled, and stays local otherwise.
The manual track publish endpoint accepts the same content_type,
language, and media_ids fields, so the composer can also drive
POST /api/v1/tracks/{id}/publish. Statuses, replies, quotes, activity
edits and track publications also accept an audio_import object
(upload_to_library, default true; fetch_metadata, default false;
library_id, default the author's lazily-created private "Uploads"
library): when enabled, each audio/* file attachment is imported into
the target library as a track after the post is saved, via
plan_audio_import/apply_audio_import (api/routes/files.py) and
services/import_.import_stored_audio_files, which reuses the file's
existing Track when it already backs one of the author's tracks and
treats external duplicates as keep_local. The import plan is validated
before the activity fans out so a bad library_id fails the request
without enqueued deliveries; per-file import failures are logged and
skipped so the post always succeeds. To keep the choice effective at
post time, the composer uploads browser files with
import_audio=false, which stores audio/* uploads as plain
StoredFiles (still audio-hashed, so the post-time import dedupes)
instead of auto-importing them. Users pick their default post format
through the status_content_type profile field (PATCH /api/v1/users/me), defaulting to text/markdown.
GET /users/{username}/objects/{object_id} in api/routes/federation.py
is the dereference endpoint for federated objects. Public tracks still
resolve through Track.federation_object_id to their Audio object; when
no track matches, an Activity is resolved by local_object_id or
source_id, scoped to the requested user (owner_user_id), so an object
is only served under its owner's namespace. Soft-deleted activities answer
with a Tombstone object (federation/activities.build_tombstone_object)
whose shape mirrors the object embedded in Delete(Tombstone) deliveries.
Live activities are only served when their visibility federates
(Visibility.federates: mentioned, followers, public) —
private/local objects were never distributed and answer 404.
federation/activities.build_activity_object produces the document: the
stored payload verbatim for payload-bearing activities (e.g. Like) —
except Create envelopes, whose embedded object is served instead because
the activity's source_id identifies the published object, not the
envelope — or a synthesized Note carrying the rendered content, the
activity_audience-derived to/cc, Mention tags, and inReplyTo.
Every post object — track_to_audio_object/track_to_note_object
serials, status and reply Notes, and payload-derived documents at serve
time — is stamped by pubby.allow_public_quotes with
interactionPolicy.canQuote granting automaticApproval to as:Public
(FEP-044f): quoting is always allowed, so Mastodon enables its Quote
action and auto-approves the quote once the QuoteAuthorization it
fetches matches. Restamping at serve time keeps objects published before
the policy existed quotable without rewriting stored payloads.
Because object URLs double as the objects' own url (a Note share's
permalink on remote servers), clients not accepting an ActivityStreams
media type are redirected to the SPA rather than served JSON: a
track-resolved object redirects to /tracks/{id}, an activity-resolved
object to /{entity_type}s/{entity_id}/activities (libraries for
library).
Served track objects are built by _track_object_document, which extends
the track_to_audio_object payload with a top-level @context and the
public to/cc audience — remote fetchers (e.g. Mastodon's URL lookup)
reject context-less documents, and an audience-less object would be
imported as a direct-only status. The object's attributedTo leads with
the publishing actor rather than the artist page URL because remote
importers take its first entry as the author and only actor URLs are
dereferenceable. GET /tracks/{track_id} content-negotiates on top of
that: requests accepting application/activity+json/application/ld+json
receive the same Audio document (the track page URL is the text/html
url advertised in every published object, so remote servers fetch it when
a user pastes the link into a search box), while browsers receive the SPA
shell annotated with a Link: rel="alternate" header and a
<link rel="alternate" type="application/activity+json"> element pointing
at the object URL. The object is only served while the track is published
(federation_object_id set). When no Audio exists, ActivityPub fetches
are redirected (303 See Other) to the track's earliest live local create
activity — _earliest_track_post picks the oldest non-deleted,
federating-visibility row and the object route then serves that share's
Note document — so a URL search on an Audio-less track still resolves to
a post (remote fetchers like Mastodon follow the redirect and import the
redirected object under its own id). Tracks with neither a published
object nor live shares answer 404 so a remote fetch cannot resurrect a
retracted post under a different id.
GET /activities/{activity_id} applies the same content negotiation to
the SPA's activity permalink — the URL ActivityCard offers for copying —
so pasting it into a remote search box resolves too. ActivityPub fetches
of a local activity receive the same document the canonical object route
serves (_activity_object_response, shared with get_object), gated on
the owner being an active local user; remote-sourced activities redirect
(303 See Other) to their origin's object id, which stays authoritative.
Browsers receive the SPA shell annotated with the Link/
<link rel="alternate"> discovery hints pointing at source_id, the
canonical object URL.
Genres
Genres are stored in a dedicated Genre table and linked to Track and
Album through GenreTrack and GenreAlbum association tables. The free-text
Track.genre and Album.genre columns remain the source of truth for embedded
metadata round-trips, while the normalised tables enable browsing, counting, and
filtering. GenreTrack.inherited distinguishes album-inherited values from
explicit track-level overrides. services/genres.py validates names, manages
associations, propagates album genres down to tracks that have no explicit
genre of their own, and re-derives the album genre from the intersection of its
tracks' explicit genres. The tag system receives the same genre-derived
tags: the genre string is split and mapped to valid tag names so
genre-derived tags appear alongside user-created ones. The public API
exposes global genre listing and deletion in api/routes/genres.py, and
per-resource genre management is supported through POST/DELETE sub-routes on
tracks and albums as well as the genre field on track/album updates. The
frontend mirrors the tag browsing experience: GenresView and GenreView
list and filter genres, GenreInput/GenreList let users edit and display
genres on tracks and albums, and the sidebar provides a top-level "Genres"
navigation link.
External Libraries
Songhive supports attaching external music libraries (e.g., cloud storage
adapters) to user- or admin-owned Songhive libraries. The feature is built
around a pluggable adapter model in songhive/external/, with a per-provider
registry (songhive/external/registry.py) and a BaseExternalLibraryAdapter
interface.
Adapters live in child plans and are registered at import time. Each adapter
exposes capabilities (list_items, read_metadata, open_stream, download,
write_tags, rename_source, delete_source, compute_hash) and a sanitized
configuration schema. Provider-specific implementations are kept outside the
core codebase and register through songhive.external.registry.
Models:
ExternalLibrary— adapter instance, encrypted provider config, capabilities, and scope (useroradmin).ExternalTrack— a track discovered through a file-oriented external provider, tied to anExternalLibraryand optionally to a SonghiveTrack.ExternalItem— an entity discovered through an entity-backed provider (kindintrack/album/artist/playlist), tying a Songhive entity to its provider-side reference (etag, mtime, raw provider metadata).ExternalSyncRun— a record of each sync attempt, including status, triggered-by, and error details.
Two provider families exist behind the same adapter interface. File
providers (local, s3, gdrive, sftp, webdav, dropbox, http)
enumerate remote audio files; each produces an ExternalTrack row that keeps
file-oriented state (sha256, mime, size, write-back/rename/delete state), and
the owning Track gets audio_file_id = NULL. Entity-backed providers
(identified by capabilities.limits["entity_import"], currently jellyfin,
tidal, and youtube) enumerate tracks, albums, artists, and playlists as provider items;
each imported Songhive entity gets an ExternalItem row in the matching kind
instead. The generic Track.source stays "external" for both — provenance is
recovered through ExternalItem.external_library.provider_type — and audio
bytes stay remote: at playback/download time the provider resolves the item
into a stream that Songhive proxies, so provider credentials never reach the
client (safe_to_redirect = False).
API surfaces:
api/routes/external_libraries.py— user-facing CRUD, sync, sync-run listing, and track management under/api/v1/external-libraries/....api/routes/admin_external_libraries.py— admin CRUD, sync, and track management under/api/v1/admin/external-libraries/....
User-created external libraries are governed by
config.external_libraries.allow_user_created_libraries and the
allow/deny provider lists. Admin libraries can be included in the public library
index when allow_admin_library_index_inclusion is enabled.
The built-in local provider (external/_local.py) lets admins mount
server-side directories as external libraries. Root paths are restricted to an
allowlist configured in external_libraries.local_roots for safety. Files under
the root are enumerated recursively, metadata is read from embedded tags, and
the directory is treated as a local audio source for streaming and download.
A separate filesystem watchdog process (external/watchdog.py) monitors enabled
local libraries and enqueues incremental or scoped syncs through
sync_external_library_task as files change, with a quiet window to coalesce
bursts of events. The watchdog is started via songhive watch-external-libraries
and is also available as a watcher container in docker-compose.yml. If no
local libraries are configured, the process polls the database every few seconds
and starts watching automatically once libraries are created, so it does not need
to be restarted after the stack is already running.
The watcher is intentionally a standalone process rather than a thread inside the
main web server: filesystem watches can be long-lived, are best as a single
instance per host, and have a different failure/restart profile than HTTP
workers. In a non-Docker deployment you should run it as a separate process
(songhive watch-external-libraries) or under a supervisor such as systemd.
Local provider settings
A local external library stores the following adapter config in ExternalLibrary.config:
| Key | Required | Default | Description |
|---|---|---|---|
root |
yes | — | Absolute or relative path to the directory to scan. |
extensions |
no | all audio | List of file extensions to index, e.g. [".mp3", ".flac"]. |
recursive |
no | true |
Whether to scan subdirectories. |
follow_symlinks |
no | false |
If false, symlinks are ignored; if true, resolved targets must stay within root. |
exclude |
no | [] |
fnmatch patterns applied to root-relative paths. |
allow_hashing |
no | true |
Whether to compute a SHA-256 hash for new/updated files. |
fast_hash |
no | false |
If true, hash the raw file bytes; otherwise use ffmpeg audio-only hashing. |
allow_write_tags |
no | false |
Whether Songhive metadata edits may be written back to the source files. |
allow_rename_source |
no | false |
Whether the rename_source operation may rename the backing file. |
allow_delete_source |
no | false |
Whether the delete_source operation may remove the backing file. |
The global external_libraries.local_roots allowlist is required: a library's
root must resolve to a path inside one of the configured roots, or validation
fails.
S3 provider
The s3 provider (external/_s3.py) indexes audio objects stored in an
S3-compatible bucket (AWS S3, MinIO, and other S3 API-compatible services).
Both regular users (when allow_user_created_libraries permits the provider)
and admins can attach buckets; like every external library, the backing
Library is private to its owner by default and can be made public/local
or shared explicitly through library share grants.
An S3 external library stores the following adapter config:
| Key | Required | Default | Description |
|---|---|---|---|
bucket |
yes | — | Bucket containing the audio objects. |
prefix |
no | "" |
Only index objects under this key prefix. |
endpoint_url |
no | AWS | Custom endpoint for S3-compatible services (e.g. http://minio:9000). |
region |
no | — | Bucket region. |
access_key/secret_key |
no | ambient | Static credentials; when omitted the SDK's ambient credentials are used. |
path_style |
no | false |
Force path-style addressing (needed by some S3-compatible services). |
presigned_urls |
no | true |
Redirect clients to short-lived presigned GET URLs for playback. |
presigned_expiry_seconds |
no | 3600 |
Presigned URL lifetime, clamped to 60–604800 (SigV4's one-week limit). |
extensions |
no | all audio | List of object-key suffixes to index. |
exclude |
no | [] |
fnmatch patterns applied to prefix-relative keys. |
recursive |
no | true |
When false, only index keys directly under prefix (delimiter /). |
allow_hashing |
no | true |
Whether to compute audio hashes for new/updated objects. |
fast_hash |
no | false |
Hash raw object bytes instead of ffmpeg audio-only hashing. |
allow_write_tags |
no | false |
Rewrite embedded tags by re-uploading the object in place. |
allow_rename_source |
no | false |
Allow rename_source (server-side copy + delete). |
allow_delete_source |
no | false |
Allow delete_source to remove objects. |
Streaming defaults to presigned GET URLs returned with safe_to_redirect
(the client fetches bytes directly from the object store, including HTTP
ranges). When presigned_urls is false, Songhive proxies the object body
through the normal stream handler, so bandwidth flows through the server.
Hashing prefers the object's S3 ChecksumSHA256 attribute when present,
then falls back to fast_hash streaming or a temp-file ffmpeg pass —
all gated on allow_hashing because they require downloading the object.
S3 freshness vs. filesystem watching
The filesystem watchdog only applies to the local provider — object stores
have no watchfiles-style change events, so S3 libraries rely on scheduled
syncs (sync_interval_seconds, driven by scan_scheduled_syncs_task).
To keep polling cheap, iter_items lists the bucket prefix with
ListObjectsV2 (paginated, no object payloads) and reports each object's
ETag, LastModified, and size. external/sync.py advertises this through
the detect_changes capability: when the stored provider_etag (or
mtime+size) still matches, the item is treated as unchanged and sync skips
metadata reads, hashing, and downloads entirely. A scheduled sync over an
unchanged bucket therefore costs one ListObjectsV2 call set — no per-object
transfer — while still catching new, modified, and deleted keys on the next
interval. Operators wanting tighter freshness should shorten
sync_interval_seconds rather than run manual full scans; scoped syncs
(scope=) can restrict a sync to a sub-prefix. Future integrations could
plug S3 event notifications or S3 Inventory into the same task pipeline.
SFTP provider
The sftp provider (external/_sftp.py, asyncssh) indexes audio files on a
remote host reached over SSH/SFTP. The remote host must be reachable from
the Songhive instance — every listing, download, and mutation originates
server-side, so firewalls/NAT between Songhive and the SSH server must allow
outbound connections to the configured host and port. Like S3, both regular
users (when allow_user_created_libraries permits the provider) and admins
can attach SFTP roots.
An SFTP external library stores the following adapter config:
| Key | Required | Default | Description |
|---|---|---|---|
host |
yes | — | Hostname or IP of the SSH server. |
port |
no | 22 |
SSH port. |
username |
yes | — | SSH login user. |
password |
no | — | SSH password (stored encrypted). |
private_key |
no | — | PEM/OpenSSH private key (stored encrypted). |
private_key_passphrase |
no | — | Passphrase for an encrypted private_key (stored encrypted). |
verify_host_key |
no | true |
Verify the server host key against known_hosts or the default files. |
known_hosts |
no | — | Inline OpenSSH known_hosts lines used when verify_host_key is on. |
root |
no | . |
Remote directory to index; relative paths resolve against the login home. |
connect_timeout |
no | 15 |
Seconds to wait for TCP connect and SSH login. |
extensions |
no | all audio | List of file extensions to index. |
exclude |
no | [] |
fnmatch patterns applied to root-relative paths. |
recursive |
no | true |
Whether to scan subdirectories. |
follow_symlinks |
no | false |
Whether to follow symbolic links while scanning. |
allow_hashing |
no | true |
Whether to compute audio hashes for new/updated files. |
fast_hash |
no | false |
Hash raw file bytes instead of ffmpeg audio-only hashing. |
allow_write_tags |
no | false |
Rewrite embedded tags by re-uploading the file in place. |
allow_rename_source |
no | false |
Allow rename_source to rename remote files. |
allow_delete_source |
no | false |
Allow delete_source to remove remote files. |
Authentication accepts a password, an inline private key (optionally
passphrase-protected), or both — asyncssh presents whichever the server
accepts. With neither configured, the Songhive process's default SSH client
keys and agent are tried, mirroring the S3 adapter's ambient-credentials
behaviour.
verify_host_key defaults on and checks the host key against the inline
known_hosts data or the process's ~/.ssh/known_hosts; disabling it must be
an explicit verify_host_key = false, since an unverified connection is open
to MITM impersonation (including credential theft when password auth is used).
SFTP has no presignable URL, so open_stream always returns a proxied byte
iterator (with byte-range support via positioned reads). Freshness follows
the S3 scheduled-sync model: iter_items walks the remote tree with
READDIR (metadata only) and reports an mtime:size change token via
detect_changes, so unchanged trees cost a directory walk and no file
transfer.
WebDAV provider
The webdav provider (external/_webdav.py, httpx) indexes audio files
stored on a remote WebDAV server (Nextcloud, ownCloud, nginx WebDAV, etc.).
The server must be reachable from the Songhive instance — every listing,
download, and mutation originates server-side, so firewalls/NAT between
Songhive and the WebDAV host must allow outbound HTTP or HTTPS connections.
Like S3 and SFTP, both regular users (when allow_user_created_libraries
permits the provider) and admins can attach WebDAV roots.
A WebDAV external library stores the following adapter config:
| Key | Required | Default | Description |
|---|---|---|---|
url |
yes | — | Base WebDAV URL, e.g. https://nextcloud.example.com/remote.php/dav. |
root |
no | "" |
Remote directory under the base URL to index. |
username |
no | — | HTTP Basic auth username. |
password |
no | — | HTTP Basic auth password (stored encrypted). |
token |
no | — | Bearer token sent in the Authorization header. |
verify_ssl |
no | true |
Verify TLS; set to false to disable or to a CA bundle path. |
timeout |
no | 30 |
HTTP request timeout in seconds. |
extensions |
no | all audio | List of file extensions to index. |
exclude |
no | [] |
fnmatch patterns applied to root-relative paths. |
recursive |
no | true |
Whether to scan subdirectories. |
allow_hashing |
no | true |
Whether to compute audio hashes for new/updated files. |
fast_hash |
no | false |
Hash raw file bytes instead of ffmpeg audio-only hashing. |
allow_write_tags |
no | false |
Rewrite embedded tags by re-uploading the file in place. |
allow_rename_source |
no | false |
Allow rename_source to rename remote files. |
allow_delete_source |
no | false |
Allow delete_source to remove remote files. |
Authentication accepts a username/password pair (HTTP Basic auth) or a
token (Bearer auth). verify_ssl defaults on and should only be disabled
for testing or trusted private networks, since disabling it leaves the
connection open to MITM impersonation and credential theft.
WebDAV has no presignable URL, so open_stream always returns a proxied byte
iterator with HTTP Range support. iter_items walks the remote tree with
WebDAV PROPFIND (metadata only) and reports an ETag/mtime/size change token
via detect_changes, so unchanged trees cost a directory walk and no file
transfer. As with S3 and SFTP, freshness comes from scheduled syncs rather
than filesystem watching.
Dropbox provider
The dropbox provider (external/_dropbox.py, httpx) indexes audio files
stored in a Dropbox account through the Dropbox HTTP API v2. All traffic
originates server-side, so the Songhive instance needs outbound HTTPS access
to api.dropboxapi.com and content.dropboxapi.com (plus
dl.dropboxusercontent.com when clients fetch temporary links directly).
Like the other providers, both regular users (when
allow_user_created_libraries permits the provider) and admins can attach
Dropbox folders.
A Dropbox external library stores the following adapter config:
| Key | Required | Default | Description |
|---|---|---|---|
access_token |
yes* | — | Dropbox OAuth access token (short-lived, ~4 hours). |
refresh_token |
yes* | — | OAuth refresh token — the durable option; mints access tokens on demand. |
app_key |
with refresh_token |
— | App key of the Dropbox app the refresh token was issued to. |
app_secret |
no | — | App secret of the Dropbox app, if it has one (stored encrypted). |
account_id |
no | — | Dropbox account id, filled in by the OAuth connect flow. |
root |
no | "" |
Dropbox folder to index (e.g. /Music); empty indexes the account root. |
timeout |
no | 30 |
HTTP request timeout in seconds. |
temporary_links |
no | false |
Redirect clients to short-lived get_temporary_link URLs for playback. |
extensions |
no | all audio | List of file extensions to index. |
exclude |
no | [] |
fnmatch patterns applied to root-relative paths. |
recursive |
no | true |
Whether to scan subfolders of the root. |
allow_hashing |
no | true |
Whether to compute audio hashes for new/updated files. |
fast_hash |
no | false |
Hash raw file bytes instead of ffmpeg audio-only hashing. |
allow_write_tags |
no | false |
Rewrite embedded tags by re-uploading the file in place. |
allow_rename_source |
no | false |
Allow rename_source to rename Dropbox files. |
allow_delete_source |
no | false |
Allow delete_source to remove Dropbox files. |
*Exactly one credential path is required: a static access_token, or a
refresh_token + app_key (+ optional app_secret) grant. Since Dropbox
access tokens only live ~4 hours, the refresh grant is the durable option —
the adapter exchanges it via oauth2/token, caches the minted access token
process-wide for its declared lifetime, and retries once with a fresh token
when a cached one is rejected.
The Dropbox app itself must be a "scoped access" app with
account_info.read, files.metadata.read and files.content.read
permissions (plus files.metadata.write and files.content.write when tag
write-back, rename or delete is wanted). Scope grants are not retroactive:
enabling a permission on the app does not extend tokens that were already
issued — after changing app permissions the account must be re-connected so
a new token is granted with the updated scope set.
Streaming defaults to proxying the file body through files/download, which
honours Range headers — it works in every browser. Setting
temporary_links redirects clients to short-lived get_temporary_link URLs
instead (safe_to_redirect; the client fetches bytes directly from Dropbox,
including HTTP ranges; links are valid ~4 hours). That offloads bandwidth to
Dropbox's CDN but fails in Firefox: dl.dropboxusercontent.com serves
Content-Security-Policy: sandbox, which Firefox applies to media loads —
enable it only for Chromium/Safari audiences. Tag write-back re-downloads the
file, rewrites tags locally, and re-uploads it in place — payloads above
Dropbox's 150 MiB single-upload cap go through chunked upload sessions
automatically.
iter_items pages files/list_folder (metadata only) and reports each
file's rev as the change token via detect_changes, so unchanged trees
cost a listing and no file transfer. The provider content_hash (SHA-256 of
4 MiB block hashes) is kept on ExternalItemRef.checksum but is never
treated as the audio sha256. As with the other remote providers, freshness
comes from scheduled syncs rather than filesystem watching.
OAuth connect flow
Providers that authenticate through OAuth (dropbox and gdrive today;
Spotify, Tidal and YouTube are planned) offer a "Connect" button in the
external-library form instead of making users paste tokens by hand. The
machinery is provider-agnostic (external/oauth.py): each provider
registers an OAuthProviderSpec describing its authorize/token endpoints,
which config keys carry the OAuth client id/secret, and how token-response
fields map onto adapter config keys. A spec may also set
scopes_for_config to derive the requested OAuth scopes from the submitted
config — Google Drive uses it to request drive.readonly or the full
drive scope depending on the form's write flags. Provider listings report
oauth_supported and oauth_callback_url (the public
{base}/api/v1/external-libraries/oauth/callback URL built from
public_base_url) so the frontend only shows the button where it applies
and can tell the user which redirect URI to register in the provider's app
console.
The flow runs in three steps:
POST /api/v1/external-libraries/oauth/begin(authenticated) validates the provider type and stores a pending entry in Redis — bound to the initiating user, the client credentials, a PKCE verifier, the callbackredirect_uriand the page to return to — under a randomstatekey (10-minute TTL). The response carries the provider's authorization URL, which the SPA navigates to after stashing the form insessionStorage.GET /api/v1/external-libraries/oauth/callback(unauthenticated; thestateparameter binds it to the pending flow) consumes the pending entry once, exchanges the authorization code for tokens server-side and stores the resulting config fragment under a result key (5-minute TTL) before redirecting back to the SPA. Provider errors and exchange failures redirect withoauth_errorinstead;return_tois restricted to local paths.POST /api/v1/external-libraries/oauth/claim(authenticated) lets the SPA collect the fragment once — results are deleted on claim and a claim by a different user leaves the entry in place.
The granted fragment is merged into the provider configuration and only
persisted when the user saves the form, at which point it goes through the
same validation and Fernet encryption as manually entered credentials.
Dropbox uses token_access_type=offline and Google Drive
access_type=offline with prompt=consent so the granted refresh_token
is durable; a future provider only needs to register a spec — no new
routes, Redis plumbing or frontend redirect handling.
Google Drive provider
The gdrive provider (external/_gdrive.py, httpx) indexes audio files
stored in a Google Drive folder, including shared (team) drives. The
Songhive instance must be able to reach googleapis.com — every listing,
download, and mutation originates server-side through the Drive API v3.
Like the other providers, both regular users (when
allow_user_created_libraries permits the provider) and admins can attach
Drive roots.
A Google Drive external library stores the following adapter config:
| Key | Required | Default | Description |
|---|---|---|---|
client_id |
OAuth | — | OAuth 2.0 client ID of a Google Cloud "Web application" client; the |
Connect flow fills access_token/refresh_token from it. |
|||
client_secret |
OAuth | — | OAuth 2.0 client secret (stored encrypted). |
refresh_token |
OAuth | — | OAuth refresh token; granted by the Connect flow (stored encrypted). |
access_token |
yes* | — | Static bearer token; expires after ~1h — testing only. |
service_account_key |
yes* | — | Service account JSON (client_email/private_key); takes precedence. |
root_folder_id |
no | root |
Folder ID to index; defaults to the My Drive root, or the shared drive |
root when drive_id is set. |
|||
drive_id |
no | — | Shared drive ID; enables supportsAllDrives and scopes lists to it. |
token_uri |
no | OAuth token endpoint override. | |
verify_ssl |
no | true |
Verify TLS; set to false to disable or to a CA bundle path. |
timeout |
no | 30 |
HTTP request timeout in seconds. |
extensions |
no | all audio | List of file extensions to index. |
exclude |
no | [] |
fnmatch patterns applied to display paths and file names. |
recursive |
no | true |
Whether to scan subfolders. |
allow_hashing |
no | true |
Whether to compute audio hashes for new/updated files. |
fast_hash |
no | false |
Hash raw file bytes instead of ffmpeg audio-only hashing. |
trash_on_delete |
no | true |
Move files to the Drive trash instead of permanently deleting. |
allow_write_tags |
no | false |
Rewrite embedded tags by re-uploading the file in place. |
allow_rename_source |
no | false |
Allow rename_source to rename remote files (file ID stays stable). |
allow_delete_source |
no | false |
Allow delete_source to trash or remove remote files. |
*Exactly one credential mode is required: OAuth user credentials, a service
account key, or a static access token. The UI "Connect" button runs the
generic OAuth connect flow against accounts.google.com — the user
registers a "Web application" OAuth client in the Google Cloud console with
the reported oauth_callback_url as an authorized redirect URI, and the
granted refresh_token (requested with access_type=offline +
prompt=consent, so re-connects always return one) lands in the config
automatically. The authorize request asks for drive.readonly, or the full
drive scope when any write flag is enabled in the form. Service account
access tokens are minted locally with an RS256 JWT bearer grant (using the
cryptography package already required for secret-at-rest encryption) and
request the same scope rule. OAuth tokens are refreshed through the token
endpoint and cached for their stated lifetime.
Unlike path-addressed providers, Google Drive items are identified by their
file ID, so provider_key carries the ID (stable across renames and moves)
and display_path carries the folder path reported by the last listing.
iter_items walks the folder tree with files.list (metadata only) and
reports the md5Checksum — a content-addressed change token — via
detect_changes, so unchanged trees cost a listing walk and no file
transfer. Google-native documents (application/vnd.google-apps.*) have no
downloadable payload and are skipped. Drive has no presignable URL, so
open_stream always returns a proxied byte iterator with HTTP Range
support; freshness comes from scheduled syncs rather than filesystem
watching, as with the other remote providers.
Jellyfin provider
The jellyfin provider (external/_jellyfin.py, httpx) is the first
entity-backed provider: instead of importing remote files, it imports a
Jellyfin server's music library as first-class Songhive tracks, albums,
artists, and playlists. Each imported entity is anchored by an ExternalItem
row (kind + provider_key + etag/mtime/raw provider metadata) instead of
an ExternalTrack, so there is no sha256, write-back, rename, or delete
state — the provider is strictly read-only, and audio bytes stay on the
Jellyfin server.
A Jellyfin external library stores the following adapter config:
| Key | Required | Default | Description |
|---|---|---|---|
server_url |
yes | — | Base URL of the Jellyfin server; SSRF-validated (public addresses only). |
api_key |
yes* | — | Jellyfin API key (Dashboard → API Keys). |
username+password |
yes* | — | Alternative to api_key; exchanged for a session token once. |
verify_ssl |
no | true |
Verify the server's TLS certificate. |
timeout |
no | 30 |
HTTP request timeout in seconds. |
collections |
no | [] |
Jellyfin library names/ids to import; empty means all music libraries. |
include_tracks |
no | true |
Import Audio items as tracks. |
include_artists |
no | true |
Import MusicArtist items as artists. |
include_albums |
no | true |
Import MusicAlbum items as albums. |
include_playlists |
no | true |
Import Playlist items as provider-owned playlists. |
sync_metadata |
no | false |
Run Songhive-side MusicBrainz enrichment on imported tracks. |
sync_cover_art |
no | true |
Reference Jellyfin image URLs as artwork (never downloaded). |
dedup_musicbrainz |
no | false |
Match artists/albums by MusicBrainz id before name matching (DB-only). |
*Exactly one auth path is required: api_key (preferred) or
username+password.
Tracks enumerate /Items?IncludeItemTypes=Audio&Recursive=true with
Fields=Etag,DateModified,Path,Genres,MediaSources,ProviderIds and an
incremental date filter when the previous run supplies since; albums,
artists, and playlists enumerate the matching IncludeItemTypes, and
playlists additionally fetch /Playlists/{id}/Items to preserve order.
Imported entities deduplicate by normalized name (or MusicBrainz id when
dedup_musicbrainz is set — a pure database lookup that never calls the
MusicBrainz API). Provider-owned albums/artists are deleted on reconciliation
only when no other provider or non-provider content still references them;
provider-owned playlists are removed with the library. raw_metadata keeps
the provider's own fields (genre list, cover URL, MusicBrainz ids, stream
Size) so entity tracks carry no local file state.
Playback resolves GET /Audio/{item_id}/stream at play time and returns the
URL with safe_to_redirect=False, so Songhive always proxies it and forwards
the client's Range header — the API key never reaches the browser.
write_tags, rename_source, delete_source, and compute_hash are
unsupported and raise UnsupportedExternalOperation.
TIDAL provider
The tidal provider (external/_tidal/, tidalapi + a throttled raw-JSON
client) imports a subscriber's TIDAL collection — saved tracks, artists,
albums, and playlists — as first-class Songhive entities. It is entity-backed
(entity_import), declares immutable_tracks (provider metadata is
write-once), lazy_contents for playlists and albums (contents are fetched on
demand, never eagerly during sync), editable_fields: ["genres", "tags"]
(everything else is provider-owned), federate_audio: false (federated
objects carry metadata and a tidal.com link only — proxied audio bytes never
leave the instance), and search.
Authentication. TIDAL credentials come from the device-authorization
framework (external/device_auth.py): the SPA calls
POST /external-libraries/device-auth/begin, the user approves a short code
on TIDAL's link page, and the SPA polls /device-auth/poll until granted;
POST /device-auth/complete then returns the credential fragment merged into
the library config. A PKCE variant (mode: "pkce", stored server-side code
verifier, user pastes the final redirect URL) is required for
HI_RES_LOSSLESS streaming — effective_quality silently downgrades to
LOSSLESS when is_pkce is false. Access/refresh tokens, is_pkce,
user_id, country_code, and expiry_time live in the encrypted library
config; sessions are cached in Redis and refreshed transparently.
Catalog cache. services/provider_catalog.py maintains
provider_catalog_entries — instance-wide raw provider payloads keyed by
(provider_type, kind, provider_key). Track entries are write-once
(expires_at = NULL): TIDAL track metadata is treated as immutable by id, so
materializing a known id never hits the API, even across users. The cache is
never exposed through the API and can be force-refreshed with
songhive admin tidal-refresh-catalog (--track ID | --album ID | --artist ID | --playlist UUID | --expired).
Lazy contents. Playlist and album items are not synchronized recursively.
ExternalItem.contents_fetched_at/contents_error track freshness; opening a
provider-backed playlist or album (GET /playlists/{id} or
/albums/{id} returns provider_sync) calls
external/lazy.ensure_contents, which enqueues refresh_external_contents
(Celery) behind a per-item Redis lock when the TTL
(playlist_ttl_seconds, instance floor 300s) has elapsed. The task fetches
the ordered track refs (etag-conditional when supported), rewrites
PlaylistTrack rows, materializes missing tracks as
membership="referenced" (full syncs reconcile by membership so referenced
tracks untouched by the collection are demoted rather than deleted), and
publishes an external_contents_refreshed WebSocket event so the frontend
reloads. A provider 404 is terminal, not retried: the item is marked
state="missing", its catalog row pinned unavailable_at, and
ensure_contents stops auto-enqueueing until a listing sync re-activates
the item. POST /playlists/{id}/provider-sync (and the album equivalent)
forces a refresh for owners/managers.
Streaming and downloads. external/_tidal/stream.py resolves playback at
request time: BTS qualities return a redirectable CDN URL (or a proxied
stream when redirect_streams is off); HI_RES_LOSSLESS manifests arrive as
MPEG-DASH — mpd_mode = "segments" serves raw segment URLs, "remux" feeds
them through external/_tidal/remux.py into the reusable
services/remote_audio_cache.py ffmpeg remux cache (bounded by the instance's
remote_cache_dir/remote_cache_max_bytes/remote_cache_retention_seconds
settings). The per-library
stream_policy limit gates every stream entry point through
services/streaming.py. Downloads (download_format = "flac" | "aac") are
disabled unless external_libraries.tidal.allow_downloads is set in the
instance config.
Metadata edits. limits["editable_fields"] is enforced by
api/routes/_common.enforce_editable_fields on the track/album/artist update,
image, tag, and genre endpoints; responses expose editable_fields and
external_provider_type so the SPA locks provider-owned inputs while keeping
tags, genres, and visibility editable.
Search. adapter.search resolves tidal.com URLs directly —
mapping.parse_tidal_url accepts both the canonical /browse/{kind}/{id}
shape and the share-link /{kind}/{id} shape (track, album, artist,
playlist) on tidal.com and its subdomains, and the entity is fetched by
id instead of text-searched. Non-TIDAL URLs and unsupported kinds
(mix, video, ...) return no results; plain text goes through the
catalog search endpoint.
A TIDAL external library stores the following adapter config:
| Key | Required | Default | Description |
|---|---|---|---|
access_token/refresh_token |
yes* | — | OAuth credentials; populated by the Connect flow. |
is_pkce |
no | false |
Set by the PKCE connect flow; gates HI_RES_LOSSLESS. |
user_id/country_code |
auto | — | Account identity captured during authorization. |
session_id/token_type/expiry_time |
auto | — | Session bookkeeping refreshed transparently. |
quality |
no | LOSSLESS |
LOW, HIGH, LOSSLESS, or HI_RES_LOSSLESS (PKCE only). |
mpd_mode |
no | segments |
segments or remux for MPEG-DASH (Hi-Res) delivery. |
redirect_streams |
no | true |
Redirect clients to TIDAL CDN URLs instead of proxying bytes. |
download_format |
no | flac |
flac or aac; requires allow_downloads in instance config. |
include_tracks/include_artists/include_albums/include_playlists |
no | true |
Which saved collections to import. |
include_followed_playlists |
no | false |
Also import followed (not owned) playlists. |
playlist_ttl_seconds |
no | 21600 |
Lazy contents cache TTL; instance floor is 300s. |
sync_metadata |
no | false |
Run MusicBrainz enrichment on imported tracks. |
dedup_isrc |
no | false |
Match imported tracks by ISRC before creating new rows. |
max_requests_per_second |
no | provider | Throttle for raw API calls. |
request_timeout_seconds |
no | provider | HTTP timeout for API calls and token refresh (default 30s). |
*Credentials are supplied by the device-auth/PKCE connect flow rather than entered manually.
YouTube provider
The youtube provider (external/_youtube/, ytmusicapi + yt-dlp +
a raw YouTube Data API v3 client) imports a Google account's liked/saved
videos, playlists, channel subscriptions and — for YouTube Music Premium
accounts — saved albums. It is entity-backed (entity_import) and shares
the TIDAL capability set: immutable_tracks, lazy_contents for playlists
and albums, editable_fields: ["genres", "tags"],
federate_audio: false (federated objects carry metadata and a
youtube.com link only), and search.
API surface. Which upstream API a library uses is resolved once per
credential set and cached in Redis for an hour: music (ytmusicapi's
innertube endpoints) for Premium accounts and browser sessions, youtube
(Data API v3 under the OAuth Bearer token) otherwise. auto mode probes
YTMusic.get_account_menu for a Premium entry; music/youtube can be
pinned per library via api_mode.
Authentication. The device-authorization flow
(external/device_auth.py) wraps Google's OAuth device grant for
TV/limited-input clients — the SPA shows the user code, the user approves
on google.com/device, and tokens (access_token/refresh_token/
expiry_time) land in the encrypted library config, refreshed
transparently through a Redis-cached, per-account locked refresh. It needs
an instance-level OAuth client of the "TVs and limited input devices" type
(external_libraries.youtube.client_id/client_secret) with the YouTube
Data API v3 enabled. Without one — or for cookie-only sessions — users
paste the request headers of a logged-in music.youtube.com browser
session into request_headers (or a Netscape export into cookies);
external/_youtube/cookies.py normalizes either shape to a Netscape cookie
file that is handed to yt-dlp on every extraction, so member/age-restricted
videos resolve like the user's own session. Both fields are credential
material and are redacted from API responses by the secret-name heuristic
(services/secrets.py also treats cookie/header keys as secret).
Streaming. There are no local bytes: at play time
external/_youtube/ytdlp.py extracts the video's format list and resolves
a signed googlevideo.com URL, cached in Redis for
stream_url_ttl_seconds (never past the URL's own expiry). The default
rendition is best audio-only (bestaudio[ext=m4a]/bestaudio/best);
GET /api/v1/stream/{id}?variant=video selects the video_format
selector (video+audio) when external_libraries.youtube.video_playback
allows it — ExternalItemRef.variant carries the choice through
services/streaming.py to open_stream. Streams stay kind="url" with
safe_to_redirect=False (googlevideo URLs are bound to the session
cookies/IP) and the stream_policy limit gates them like TIDAL. The SPA
exposes the video rendition through the reusable
components/player/VideoEmbedPlayer.vue on the track page's "Watch video"
action. Downloads (download_format = "audio" | "video") run yt-dlp into
the stream temp dir and are disabled unless
external_libraries.youtube.allow_downloads is set.
Search. adapter.search resolves youtube.com/youtu.be URLs (and
bare video ids) directly — falling back to a yt-dlp extraction when the
active API surface can't address the entity — and forwards other text to
ytmusicapi's catalog search.
A YouTube external library stores the following adapter config:
| Key | Required | Default | Description |
|---|---|---|---|
access_token/refresh_token/expiry_time |
yes* | — | OAuth credentials; populated by the Connect flow. |
client_id/client_secret |
auto | — | OAuth client pair that issued the tokens (recorded at grant). |
request_headers/cookies |
alt* | — | Pasted browser session (headers text or Netscape export). |
auth_mode |
auto | — | oauth or browser; inferred from stored credentials. |
account_name/channel_handle |
auto | — | Account identity captured after authorization. |
api_mode |
no | auto |
auto, music, or youtube. |
include_tracks/include_playlists/include_subscriptions/include_albums |
no | true |
Which collections to import (albums need music mode). |
download_format |
no | audio |
audio or video; requires allow_downloads in instance config. |
playlist_ttl_seconds |
no | 21600 |
Lazy contents cache TTL; instance floor is 300s. |
video_format |
no | 720p default | yt-dlp selector for the video+audio rendition. |
sync_metadata |
no | false |
Run MusicBrainz enrichment on imported tracks. |
max_requests_per_second |
no | provider | Throttle for YouTube API calls. |
request_timeout_seconds |
no | provider | HTTP timeout for API calls and token refresh (default 30s). |
*Either OAuth credentials or a pasted browser session is required; the device-auth flow supplies the former.
Visibility, sharing, and secret redaction
Every external library is backed by a normal Library row, so visibility
(private/local/public) and ShareGrant sharing behave exactly like
regular libraries. PATCH /api/v1/external-libraries/{id} (and the admin
equivalent) accepts a visibility field; changing it updates the backing
library and calls services.music.propagate_external_library_visibility,
which rewrites the visibility of all linked synced tracks so list queries
and single-track ACLs stay consistent. PATCH /api/v1/libraries/{id} on an
external-backed library propagates the same way. Non-admin propagation only
touches tracks the caller owns, mirroring album visibility propagation.
Because API responses redact secret-bearing config keys to "<redacted>",
both PATCH routes run submitted configs through
_merge_config_preserving_redacted: any key still carrying the sentinel is
restored from the stored (decrypted) config before validation and
re-encryption, so editing non-secret fields never clobbers credentials.
The web UI renders a per-provider configuration form instead of raw JSON.
Provider form templates live in
frontend/src/config/externalLibraryProviderTemplates.ts. Each template entry
lists the JSON property name, i18n label/description keys, field type (string,
password, number, boolean, enum, comma-separated string array, or multiline
textarea), and default value. A field may also declare a configKey to
read/write a different JSON property than its own name — used for union-typed
keys such as WebDAV's verify_ssl, where the optional CA-bundle text field
shares the key with the verification toggle and overrides it when non-empty.
ExternalLibraryEditView switches the displayed fields whenever the provider
<select> changes, and falls back to a plain JSON textarea for providers that do
not have a template yet. Template forms only emit their declared fields, so on
update the view also echoes back every stored config key the template doesn't
manage (carriedStoredConfig) — device-auth/OAuth credential fragments such as
access_token, user_id or auth_mode would otherwise be dropped by the
PATCH's full-config replacement and the save would fail validation.
Sync and tasks:
tasks/external_libraries.pydefines three Celery tasks:sync_external_library_taskruns a manual or scheduled sync, enumerating provider items throughexternal/sync.pyand upsertingExternalTrackrows for file providers orExternalItem-backed entities for entity-backed providers.scan_scheduled_syncs_taskscans for due external libraries and enqueues sync tasks while skipping libraries with an active run.write_back_metadata_taskapplies local metadata edits to the provider when the adapter supportswrite_tags.rename_sourcelets track owners and admins rename the backing source file from the track edit page when the adapter supports it.- Destructive source deletion requires both
write_tagsanddelete_sourcecapabilities, the globalallow_destructive_deletesetting, and an explicitconfirm: "DELETE"from the caller. Tracks may optionally be removed from Songhive at the same time.
Streaming and downloads:
services/streaming.resolve_external_streamloads an external track (ExternalTrackfor file providers,ExternalItemfor entity-backed providers) and opens a provider stream (path,iterator, orurl).streaming/handler.pyfalls back from localStoredFileplayback to external streams, preserving auth, ACL, history, and now-playing broadcasts.api/routes/tracks.pyprovidesGET /api/v1/tracks/{track_id}/download, proxying external bytes through Songhive and honoringContent-Dispositionanddisposition=inline|attachment.services/streaming.collect_external_streammaterializes iterator/url streams into memory or a temp file. Audio-typed payloads in MP4 containers get theirftypmajor brand rewritten toM4A— audio-only MP4s often carry a genericisom/mp4*brand that content sniffers (libmagic, and therefore Fediverse media fetchers such as Mastodon) report asvideo/mp4regardless of the servedContent-Type, which remote servers then process (and reject) as video. Non-temporarypathstreams are never rewritten — they reference files owned by the adapter (e.g. the local library). The same caveat applies to localStoredFiledownloads, which are served straight from storage and not normalized.
Secrets and audit:
services/secrets.pyencrypts/decrypts provider configs at rest and redacts them in API responses and audit logs.- All mutating external-library and external-track endpoints log an
AuditLogrow with a sanitizeddetailsdictionary.
Duplicate detection
When an uploaded audio file's audio-only SHA-256 hash matches an ExternalTrack
in active or shadowed state, services/import_.py raises an
ExternalDuplicateError. The POST /api/v1/files/upload and
POST /api/v1/files/upload/bulk endpoints return a 409 response with an
ExternalDuplicateWarning containing a short-lived Redis token.
The caller can then resolve the conflict via
POST /api/v1/files/upload/resolve-duplicate with action=keep_local (create a
local Songhive track and keep the uploaded file) or action=discard_upload
(reuse the existing external track and remove the uploaded stored file).
API Adapters (Subsonic)
Besides the native /api/v1 surface, Songhive can expose foreign media APIs
so third-party clients can browse and stream its content. These integrations
live in songhive/adapters/ — a different adapter family from
songhive/external/ (which imports remote libraries into Songhive). API
adapters map Songhive's services and ACLs onto an external protocol; the
first implementation is Subsonic/OpenSubsonic, and the registry is designed
for future Icecast/Mopidy/Jellyfin adapters.
Framework (songhive/adapters/):
base.py—APIAdapterinterface:name,is_enabled(config),router()(FastAPI routes),tornado_routes()(native handlers for streaming),include(app, config).registry.py— name-keyed registry (register_adapter,get_adapter,list_adapters),mount_adapters(app, config)invoked fromcreate_app, andadapter_tornado_routes(config)consumed by_build_tornado_appahead of the WSGI fallback.- Each adapter owns its complete URL namespace (Subsonic owns
/rest).
Subsonic adapter (songhive/adapters/subsonic/):
- Namespace —
/rest/*.view, the path Subsonic clients hard-code.ping,getLicense,getMusicFolders, browsing (getIndexes,getArtists,getArtist,getAlbum,getSong,getMusicDirectory,getArtistInfo*,getAlbumInfo*),getAlbumList*,getGenres,getSongsByGenre,getRandomSongs,search*variants, playlists (getPlaylists,getPlaylist,createPlaylist,updatePlaylist,deletePlaylist), starring (getStarred*,star,unstar),scrobble,nowPlaying,getNowPlaying,stream,download,getCoverArt,getAvatar, plus an OpenSubsonicgetOpenSubsonicExtensionsadvertisement and a/rest/{method}.viewcatch-all that returns a protocol error for unimplemented methods. - Envelope — every endpoint answers with the
subsonic-responseenvelope (protocol1.16.1,type=songhive,openSubsonic=true) in XML by default, JSON viaf=json, or JSONP viaf=jsonp+callback. JSONP callbacks are restricted to safe identifier paths to prevent script injection. - Authentication — per-request
u+p(plain,enc:-hex, or a Songhive API token inp),u+apiKey(OpenSubsonic), andu+t/ssalted-token auth. Salted tokens (md5(password + salt)) can only be verified against credentials the server can reconstruct, so the real password isp-only — but API tokens work: HS256 JWTs are deterministic given their claims, and every claim is recoverable from theapi_tokensrow (user_id,jti,expires_at, pluscreated_atwhichissue_api_tokenpins to theiatclaim).auth.pyrebuilds each candidate JWT byte-for-byte (tryingcreated_at±1s for tokens issued before the pin) and comparesmd5(jwt + s)tot, so salted-token clients work with an API token as the password while no usable credential is stored. Because credentials ride in request parameters rather than cookies,/rest/is exempt from the CSRF middleware. - ACL mapping — every browse/stream path goes through the same
services.aclchecks as/api/v1: anonymous users see public content, authenticated users see public/local/owned/shared content, inaccessible items map to error 50 (or 70 where the spec calls for not-found). - Streaming — under the Tornado bootstrap
/rest/stream.viewand/rest/download.vieware served natively byhandlers.py(subclasses ofStreamHandler, reusing its range, transcoding and external-stream logic with Subsonic auth/error mapping).media.pyregisters the same endpoints on the FastAPI router as the uvicorn fallback;getCoverArt/getAvatarare FastAPI-only. - Playback reporting —
scrobblerecords listens viaservices.streaming.record_listen;nowPlaying/getNowPlayingshare an in-processnow_playing.pyregistry (also populated when tracks are streamed). Mutating playlist endpoints writeAuditLogrows throughservices.audit.log_action. - Config —
subsonic.enabled(defaulttrue) gates mounting; nginx proxies/rest/withproxy_buffering offfor streaming.
Database Migrations
Schema changes are managed with Alembic. The
songhive/migrations/ package contains the Alembic environment, the
versions/ directory, and an empty base revision that marks the pre-
migration schema baseline. New installs receive the current schema from
Base.metadata.create_all and are then stamped at head; existing
production databases are stamped at base and upgraded normally.
Migrations run automatically when create_app is called, and the Docker
entrypoint runs songhive admin migrate before starting the web server or
Celery workers. ensure_migrated acquires a backend-specific lock for the
whole operation: a PostgreSQL advisory lock for Postgres, or a fcntl file
lock on a companion file for SQLite. This prevents the web server and worker
from racing to create the baseline schema on a fresh Docker Compose install.
Admins can also trigger them manually:
python -m songhive admin migrate
To autogenerate a revision after a model change (from the repository root,
with SONGHIVE_DATABASE__URL set or database.url configured):
alembic revision --autogenerate -m "add example column"
Authentication & Authorization
Authentication flow
- Login — username/password → bcrypt verify → issue
TokenPair(short-lived JWT access token + long-lived opaque refresh token stored in Redis). Each refresh token also records IP, user agent, creation time and expiry, forming a user session. The JSON response still carries the token pair, but the server also sets three cookies (api/cookies.py):access_tokenandrefresh_tokenareHttpOnlyso browser JavaScript never reads the credentials, while a readablecsrf_tokenbacks the double-submit CSRF check. The refresh cookie is scoped toPath=/api/v1/auth; cookie flags are controlled byauth.cookie_secure(defaults to secure outside debug mode),auth.cookie_samesite(lax), andauth.cookie_domain. - Refresh — opaque refresh token from the JSON body (API clients) or the
refresh_tokencookie (browsers) → Redis lookup → rotate (revoke old, issue new pair) and re-set the cookies. Failed refreshes clear the cookies. - Revoke — single token or all tokens for a user (Redis key deletion).
Access tokens now carry a
jticlaim, and revocation also adds that JTI to a Redis deny-list. The JWT middleware checks the deny-list so revoked sessions cannot continue using their existing access token until expiry. - Session management — users can list their active refresh-token sessions (including the current one) and revoke any session individually, revoking the current session ends the user's login.
- OAuth2 — authlib
authorization_codeflow for third-party app access. - Invite-only registration — controlled by
RegistrationModeconfig setting;Invitecodes with optionalmax_usesandexpires_at.
Authorization
- JWT middleware —
api/middleware/auth.pydecodes the bearer token (Authorization: Bearer …first, then theaccess_tokencookie) and injects the currentUservia FastAPI's dependency system. Cookie auth is what lets<img>/<audio>elements, file downloads, and the WebSocket handshake authenticate without JavaScript-accessible tokens. - CSRF —
api/middleware/csrf.pyenforces a double-submit check: unsafe methods (POST/PUT/PATCH/DELETE) that carry an auth cookie and noAuthorizationheader must echo the readablecsrf_tokencookie inX-CSRF-Token. Session endpoints (login/refresh/logout/registration, OAuth token endpoints) are exempt; bearer clients and safe methods are unaffected. - Role-based —
UserRole.ADMIN/MODERATOR/USER;require_admindependency enforces admin-only routes. - ACL service (
services/acl.py) — three-level visibility check augmented withShareGrant(owner grants a named user) andShareToken(revocable short-link cookie).Artistrows carry novisibility/owner_idof their own, so artist listings (GET /api/v1/artists/and the artists section ofGET /api/v1/search/) instead require at least one track or album the requester can access; artists whose content is entirely private or unshared are hidden. Admins bypass the filter.GET /api/v1/artists/also acceptsowner_username(like the album/track/library/playlist listings) to keep only artists with at least one track or album owned by that user, on top of the requester's ACL.StoredFilerows inherit access from the entities referencing them: a file attached to a track, album, playlist, or library is downloadable by anyone who can access that entity (directly, via aShareGrant, or via aShareToken). Album covers are also reachable through the album's tracks, and artist images through the artist's tracks and albums, so sharing a track reveals its album cover and artist image. - Rate limiting — Redis sliding-window;
rate_limit(IP),rate_limit_user_or_ip(authenticated users keyed by id), andrate_limit_account(always per-user) FastAPI dependencies. MediaDELETEendpoints userate_limit_accountfor per-user rate limiting on destructive operations. Byte-serving endpoints (GET /api/v1/files/{id}/download,GET /api/v1/tracks/{id}/download) userate_limit_file_download/rate_limit_media, which draw from the looserauth.rate_limit_media_requestsbudget (0 disables it) instead ofrate_limit_requests;image/*stored files are exempt entirely so avatars, covers and image attachments survive federation-driven fetch bursts from a single remote egress IP. Fails open when Redis is unavailable. Downloaded files are content-addressed, so their responses carryCache-Control: …, max-age=31536000, immutable—publicfor publicly visible files,privateotherwise. - CORS — the API uses the credentialed
CORSMiddlewareallowlist fromserver.cors_origins. Read-only media endpoints (GET /api/v1/files/{id}/download,GET /api/v1/tracks/{id}/download, andGET /api/v1/stream/{id}) instead returnAccess-Control-Allow-Origin: *viaapi/middleware/media_cors.py(andset_default_headersin the TornadoStreamHandler) so remote Fediverse clients can embed audio directly — this is safe because the wildcard can never be combined with credentials, and theaccess_tokencookie isSameSite=Laxso it is not sent cross-origin. Preflights on those paths allow theRangeheader and exposeContent-Range/Accept-Ranges/Content-Lengthto fetch-based players.
Owner exposure
Entity detail responses for Track, Album, Library, and Playlist include owner_id whenever the requester has ACL access to the entity; ?include=owner adds a nested UserSummary with actor_url, username, display_name, and avatar_url. The frontend useEntityMeta composable returns this full owner object, and UserLink routes local users to /@{username} and remote users to their actor_url.
Profile visibility
Each user carries a profile_visibility preference (public, local, or
private; default public, editable from /settings via
PATCH /api/v1/users/me) controlling whether their profile appears in the
users directory. GET /api/v1/users and the users section of
GET /api/v1/search/ hide local profiles from anonymous callers and never
list private ones — not even to their owner. Individual profile pages
(GET /api/v1/users/{username}, /@{username}) remain reachable regardless
of the setting.
File Storage & Upload Pipeline
- Client lists visible files via
GET /api/v1/files/, uploads viaPOST /api/v1/files/upload, or bulk-uploads viaPOST /api/v1/files/upload/bulk. The bulk endpoint uses the same per-IP rate limit as the single-file upload and enforces per-request limits on the number of files and total request size. StorageService(facade overStorageBackend) validates size limit, computes SHA-256, deduplicates by hash, writes to backend.- A
StoredFilerow is created (content-addressable, owner/visibility set). - An
Uploadrow is created linking the stored file to a track (or pending import). - A Celery
import_task processes metadata, creates/updatesTrack,Album,Artistrecords, and optionally triggersenrich_track. enrich_trackenqueuessync_track_tagsandenrich_imageson success; the latter fetches artist images and any missing album covers.- A separate Celery
transcodingtask pre-transcodes to requested formats and stores results asTranscodedFilerows pointing back to aStoredFile.
Orphan GC — the storage.cleanup_orphaned_files Celery task runs on the
configured crontab (default: daily at 03:00) and deletes StoredFile rows
(and their backing files) not referenced by any Track, Album, Upload,
or DownloadArchive (archive-backed ZIPs are protected until the archive
row itself is cleaned up).
Cascade Deletion — services/deletion.py provides centralized deletion
logic for Track, StoredFile, Album, Artist, Playlist, and Library.
DELETE endpoints on the corresponding routes accept a recursive query
parameter; albums default to recursive deletion while other collections default
to non-recursive. A dedicated DELETE /api/v1/tracks/bulk endpoint accepts a
list of track IDs and delegates to delete_tracks_bulk in the deletion service,
applying the same ACL and rate-limiting checks as single-track deletion.
Deleting a track removes its Upload, LibraryTrack,
PlaylistTrack, Favorite, ListeningHistory, TranscodedFile, ShareGrant,
ShareToken, and Report rows and deletes the underlying StoredFile once it
is unreferenced. Deleting a stored file removes all tracks that use it as their
audio source and clears cover_file_id/image_file_id references on
albums/artists before removing the backing object. Deleting an external library
cascades through its linked ExternalTrack rows, removing each Songhive track
and its dependents (playlist entries, favorites, etc.), cleaning up empty
albums/artists, and removing the underlying Library only when it is empty.
Recursive deletion collects unpublish information for public tracks and enqueues
Delete(Tombstone) ActivityPub activities.
Streaming & Transcoding
- Tornado
StreamHandler(streaming/handler.py) handlesGET /api/v1/stream/{track_id}, resolves the bestStoredFileviaservices/streaming.py, and streams with native range-request support. - Transcode cache — on first request for a (format, bitrate) combination,
the
Transcoderwrites the output to a temp file, which is then stored as aStoredFileand indexed inTranscodedFile. Subsequent requests serve the cached file directly (skipping ffmpeg) whentranscode_cache_enabled. - Bitrate enforcement —
effective_bitrate()inconfig/schema.pycomputes the minimum of: requested bitrate, instance max, and per-role max (max_bitrate_by_rolekeyed byUser.role). - Supported formats: MP3, OGG (Vorbis), FLAC, AAC (M4A), Opus.
Server-side audio outputs
Playback can be routed to persistent server-side outputs instead of the browser Web Player, Spotify-Connect-style: the browser (or any client) sends commands, but audio is rendered by the server, so playback continues after the controlling tab closes.
- Control plane — a persisted
PlaybackSessionrow (models/playback_session.py) holds the queue, index, position, volume and play/pause state per user.services/playback.pyvalidates commands, commits the new state, then publishes a control envelope on the Redis listsonghive:playback:control:{session_id}. Session changes are broadcast over/ws/events(playback_sessionevents) so multiple tabs stay in sync. - Outputs — an
OutputStreamrow (models/output_stream.py) stores aprovider_typeplus a provider config. Secret config fields are Fernet-encrypted at rest and redacted to"<redacted>"in API responses; PATCH routes merge the sentinel back to the stored value. Providers are registered instreams/registry.pyby subclassingAudioOutput(provider_type,FIELDSschema hint,validate_config,create_driver);GET /api/v1/outputs/providersexposes the schema to the frontend form builder. - Stream worker —
songhive stream-worker(streams/worker.py) is a dedicated process that scans for active sessions bound to outputs and runs oneSessionDriverper session. Each driver holds a Redis locksonghive:stream:lock:{output_id}(unique token, refreshed every loop, released on shutdown) so exactly one worker drives a given output. The worker owns queue advancement, repeat/shuffle, listen recording, scrobbling and idle shutdown (streams.background_idle_timeout_seconds). A failed track advance (transient DB/storage/provider error) is retried a few times before the driver falls back to its silence generator, and the advance is idempotent — a retry after a committed index update re-syncs the source instead of skipping a track. Deployed as thestream-workerCompose service (streamsprofile) andsonghive-stream-worker.service. - Icecast provider (
streams/icecast.py) — one long-lived ffmpeg encoder pushes an MP3, Ogg Vorbis or Opus stream to anicecast://mountpoint; a short-lived ffmpeg decoder (-re -ss <pos>) feeds PCM into it per track. Pausing swaps the decoder for a realtime (-re)anullsrcsilence generator so the mount stays alive without flooding listeners.source_endeddecoder events carry a generation tag so events from a killed decoder can't pause the new source. ffmpeg's Icecast muxer cannot update ICY metadata in flight, soupdate_metadatapublishes the now-playing title out-of-band to the server's/admin/metadata?mount=...&mode=updinfo&song=...endpoint using the source credentials (best-effort: failures are logged, never raised), and the push is repeated after an encoder reconnect because a source reconnect clears the mount's metadata. - Native HTTP provider (
streams/http.py) — same driver machinery, but the encoder writes topipe:1and a publish task XADDs each ~16 KiB chunk (base64) into the Redis streamsonghive:stream:data:{mount}(MAXLEN ≈ streams.http_stream_max_entries). The TornadoStreamMountHandler(streaming/mount.py) servesGET /streams/{mount}by bursting the neweststreams.http_stream_burst_entriesentries (XREVRANGE) then following the stream (XREAD BLOCK) — every listener is an independent cursor, so Redis provides the fan-out and a stalled client never slows the others. Entries older thanstreams.http_stream_max_lag_seconds(entry IDs are server millisecond timestamps) are skipped rather than delivered, so a listener that falls behind jumps forward instead of accumulating latency, and theX-Accel-Buffering: noresponse header keeps buffering proxies from hiding listener lag in their own buffers. Liveness is the TTL'dsonghive:stream:meta:{mount}key, refreshed by the driver on a timer so decoder gaps at track boundaries cannot expire it;{"end": "1"}entries disconnect listeners on graceful stop and on encoder restart (a fresh encoder produces a new container stream that cannot be spliced into an open response); per-listener TTL keys undersonghive:stream:listener:{mount}:*, refreshed by the handler while connected, feedlistener_countfor idle shutdown and enforcestreams.http_stream_max_listeners(0 = uncapped). Now-playing metadata rides the same stream as{"m": ...}entries and is mirrored into the meta blob; listeners that sendIcy-MetaData: 1get anicy-metaintheader andStreamTitleblocks interleaved into the audio everystreams.http_stream_metaint_bytesbytes, so track changes update the title in audio order. Track changes, driver starts and stops also emit astream_updateevent ({mount, online, now_playing}) on the WebSocket fan-out channel so the/streamsdirectory page updates live — public mounts broadcast on thestreamstopic, whilelisten_tokenmounts are private and the event is delivered only to the owner's connections (the worker injectsOutputStream.user_idinto the driver config as_owner_user_id). The page additionally pollsGET /api/v1/streams/every 30 s, both as the update path for anonymous visitors (the WS endpoint requires authentication) and to reconcile states no event can carry — a crashed driver only expires its meta key, and listener counts change constantly. Mount slugs must be unique acrosshttpoutputs; an optionallisten_tokenfield gates listeners via?token=/Bearer. No external server or extra port is needed. - Snapcast provider (
streams/snapcast.py) — same driver machinery, but the "encoder" is an ffmpeg passthrough: raw s16le stereo PCM from stdin is copied unchanged to a snapserver sink — either the named pipe of a snapserverpipe://source (mode=fifo, auto-created withmkfifo, existing non-FIFO paths rejected so a typo can't clobber a regular file) or a snapservertcp://listening source (mode=tcp), which also allows a remote snapserver. Snapcast streams are never registered dynamically: the source must already be declared insnapserver.conf(e.g.stream = tcp://0.0.0.0:4953?name=...&sampleformat=...), the targetportis that source's listener — not the snapclient port 1704 — andsample_ratemust match the source'ssampleformat.listener_countqueries snapserver's JSON-RPC control API (Server.GetStatusover raw TCP, default port 1705 — the HTTP/snapweb port 1780 does not speak the newline-delimited protocol;control_host/control_portfields) and counts connected, unmuted clients — optionally only those in groups playingstream_name— so idle shutdown tracks real listeners like the Icecast status endpoint does.update_metadatais a no-op: snapserver's control API offers no metadata setter — stream metadata can only originate from acontrolscriptplugin attached to the source insnapserver.conf. - Driver interface (
streams/driver.py) —start,stop,set_source,pause,resume,seek,set_volume,update_metadata,health,is_paused,listener_count, and aneventsqueue. The worker usesdriver.is_paused(not the persisted session state, which is already post-command by the time the worker reads it) to detect pause→play transitions.
Bulk Downloads
Authenticated users can request ZIP archives of tracks — a selection made in
the library's bulk-edit mode, or a whole album, artist, playlist, or
library — which are built asynchronously by a Celery task and downloaded
from the /downloads page when ready.
Request flow (api/routes/downloads.py,
services/downloads.py):
POST /api/v1/downloads/accepts explicit references (track_ids,remote_object_ids,episode_ids) and/or containers (album_id,artist_id,playlist_id,library_id) plus an optionallabel.services.downloads.resolve_archive_itemsexpands containers through the regular ACL — items the requester cannot access are silently skipped — and caps the result atdownloads.max_items(default 500). A request that resolves to zero accessible items is rejected.- A
DownloadArchiverow (models/download.py) is created inpendingstate with a JSONitemssnapshot (kind, id, title, artist, source hints) so the build is reproducible even if the library changes.downloads.max_active_per_user(default 2) limits how many pending/processing archives one user may hold. Creation, file download, deletion, and clear go through the shared Redis sliding-window rate limiter; the read-only list/detail endpoints are unthrottled so the page can poll freely. tasks/downloads.py'sbuild_download_archiveCelery task materializes each item into a ZIP on local disk, stores it throughStorageServiceas a privateStoredFile, and flips the row toready(orfailedwhen no item succeeded). Per-item failures are collected intoitem_errors, so a partially failed archive is still delivered. Instance-wide concurrency is capped atdownloads.max_concurrent_archives(default 2) by a Redis sorted-set semaphore (songhive:downloads:archive-slots) with expiring leases — tasks that can't take a slot requeue themselves; the limiter fails open when Redis is unavailable.- The requester gets a
downloadnotification (archive_id+labelpayload) pointing at the Downloads page.
Item materialization covers every playable source:
- Local tracks stream from their
StoredFilevia the storage backend. - External-library tracks go through the provider adapter's
download/open_stream(services/streaming.py's external-stream helpers). - Remote (federated) objects and podcast episode enclosures are fetched to
disk by
federation/fetch.py'sguarded_download— the same SSRF protection asguarded_fetch(DNS/IP checks, per-redirect-hop revalidation, remote-domain moderation), streamed with afetch_timeout_secondsbudget and amax_item_bytescap (default 2 GiB). - Transient remote/external fetches retry
fetch_attemptstimes (default 3) withfetch_backoff_secondsexponential backoff.
ZIP member names are sanitized (01 - Artist - Title.mp3,
collision-deduped) so archives unzip cleanly on any platform.
Lifecycle — GET /api/v1/downloads/ lists the caller's archives
(newest first); GET .../{archive_id}/file streams a ready ZIP;
DELETE .../{archive_id} removes the row and its backing file;
POST .../clear removes all ready/failed rows. The hourly
cleanup_completed_downloads beat task deletes ready/failed archives
older than downloads.retention_hours (default 24) and marks
pending/processing rows stale for more than downloads.stale_run_hours
(default 6) as failed.
Frontend — TrackList.vue's bulk-edit bar, AlbumView, ArtistView
and PlaylistView expose a Download action through the shared
useDownloadArchive composable (toast feedback + link to the Downloads
page). DownloadsView (/downloads, own nav entry) polls every five
seconds while archives are active and offers per-archive download/delete
plus a confirmed "clear completed" action. api/downloads.ts and the
regenerated types.ts carry the typed client.
Payments
Optional payments support (payments.enabled, [payments] config) covers
two flows: artist sales (fans buy tracks or albums; sellers are paid
through their own Stripe Connect accounts) and paid instance
memberships (registration_mode = "paid" gates accounts behind an
instance subscription, default $5/month configurable via admin settings).
Provider abstraction (services/payments/providers/) — a small
PaymentProvider protocol with a Stripe implementation (Checkout Sessions
with direct charges on connected accounts for sales, Billing for
memberships, Connect for seller onboarding) and a deterministic fake
provider used by the test suite. Seller accounts are created through the
Accounts v2 API (POST /v2/core/accounts) — Stripe disables Accounts v1
creation for new Connect integrations — with merchant.card_payments
(direct charges) and recipient.stripe_transfers capabilities,
identity.country taken from the seller's onboarding request (required
before a merchant configuration can be applied). payments.stripe_dashboard
picks the hosted dashboard and pins the only defaults.responsibilities
combination Stripe accepts for it without approval: full (default) gives
sellers the standard Stripe dashboard with
fees_collector/losses_collector = stripe; express gives the
co-branded Express dashboard with platform-collected fees and
losses_collector from payments.stripe_losses_collector — application
requires managed-risk approval, stripe requires the gated Express +
Stripe-managed-liability preview (account_controller_unsupported_configuration
otherwise). Onboarding uses v2 Account Links and connected_account_state
normalizes both v1 (account.updated webhook) and v2 (live retrieve)
payload shapes into the stored ConnectedAccount flags. Provider events are the only authority that moves
money state: checkout redirects are never trusted.
Data model (models/payments.py) — Sale (per track or album, with
unpaid_policy of full_stream/sample/none), PaymentOrder +
PaymentOrderItem (both purchase and membership kinds; guest emails
are Fernet-encrypted), PurchaseEntitlement (permanent per-track rights
keyed on user_id or a guest email_hash), RedeemCapability (opaque,
expiring, rate-limited guest download tokens), ConnectedAccount,
InstanceSubscription, PurchaseArtifact/SampleDerivative (paid
snapshot ZIPs and trimmed samples as private StoredFiles), plus
PaymentEvent inbox and FulfillmentOutbox rows for idempotent webhook
processing and durable guest delivery.
Access policy (services/payments/access.py) is the single source of
truth for byte-level access: owners/sellers/admins and active entitlements
get full; unpaid listeners get the sale's policy level (stream,
sample, or none). Refunded or disputed orders revoke entitlements.
Every media path consults it — serialized track responses rewrite
audio_url/can_download, the Tornado stream handler serves original
bytes, the configured sample, or 403 by level (SubsonicDownloadHandler
flips _requires_download_access so download.view needs full), and
files/{id}/download, tracks/{id}/download, download-archive
materialization, share pages, feeds, M3U and federation serializers all
suppress or substitute URLs the requester isn't entitled to.
Membership lifecycle (services/payments/membership.py) —
sync_user_active_flag is the only writer of User.is_active for payment
reasons, combining payments_required, admin_suspended, subscription
paid_through + membership_grace_hours, and email verification. Paid
registrations start inactive; invoice.paid activates only when email is
verified (verification itself re-syncs). Unpaid-but-verified users get a
402 with a scoped Redis billing capability (billing_capability_ttl_seconds)
that authorizes only the membership endpoints. POST /payments/membership/session exchanges a billing capability or a fresh
order checkout token for a real session once the webhook-confirmed state
shows the account active — each scoped credential mints exactly one session
(Redis single-use marker) and the order token only within the capability TTL
of the order's last update, the same freshness window membership portal and
cancellation enforce for order-token callers. Admin deactivation/deletion
cancels provider subscriptions first, with durable retry intent when the
provider call can't be confirmed.
Webhooks (POST /api/v1/payments/webhooks/stripe{,-connect}, CSRF-exempt
server-to-server calls) verify the provider signature, persist a normalized
PaymentEvent inbox row, and process idempotently through
tasks/payments.py's process_payment_event. Beat-scheduled sweeps bound
lost deliveries: membership expirations, pending-order expiry, outbox
retries, and artifact cleanup.
Frontend — PurchasePanel/SaleEditor/MembershipPanel components,
/checkout/*, /redeem, /billing routes, profile Purchases/Billing tabs
(shown only when payments are enabled), login 402 → /billing?bcap=…
routing, and admin payments_required toggles on the Users page.
Federation
Federation is powered by pubby mounted
on the FastAPI app via its FastAPI adapter. Per-user actor routes and WebFinger
discovery are in api/routes/federation.py. The general-purpose primitives are
delegated to pubby:
- Outbound plaintext→HTML rendering (escaping, URL linkification,
rel="tag"/rel="me"anchors,HashtagandPropertyValuetag/attachment builders) andAudioobject content/duration formatting (set_object_content,format_duration) come frompubby.content; the instance's/tags/{name}route convention is injected viafederation/_common.py'sget_tag_url. - Instance allow/block matching (
normalize_domain,extract_domain,is_domain_blocked) comes frompubby.moderation;services/federation.pykeeps thin wrappers that injectconfig.federation.allowed_instances/blocked_instances. - Async→sync database URL conversion for pubby's synchronous SQLAlchemy
storage is
pubby.storage.adapters.db.to_sync_url(applied byinit_db_storageinsidecreate_activitypub_storage). - Actor private-key provisioning is
pubby.crypto.ensure_private_key_file. - Follower inbox collection is
pubby.collect_inboxes; one-shot signed delivery ispubby.deliver_activity(the Celery task keeps the retry policy).
Songhive keeps orchestration: Celery tasks and their retry policy, per-user
actor documents built from the User model, federation_* table naming, and
the HTTP routes.
Actor model:
- Each
Userhasactor_url,private_key_pem, andpublic_key_pemcolumns. - New users are provisioned with RSA keypairs on creation when federation is
enabled. Existing users can be back-filled via
songhive admin provision-federation-keys. - Each user is reachable at:
https://{instance_domain}/users/{username}— canonicalPersonactorhttps://{instance_domain}/@{username}— Mastodon-style aliashttps://{instance_domain}/.well-known/webfinger?resource=acct:{username}@{instance_domain}
Activity lifecycle:
- Uploaded tracks are published as
Audioobjects viaCreateactivity — only when explicitly requested. Uploads are local-only by default; the/api/v1/files/uploadendpoints (single and bulk) and the/{library_id}/tracksupload endpoints take an opt-inpublish=trueflag (surfaced in the UI as a "Publish on the Fediverse" checkbox shown when the instance federates), which the synchronous library bulk path and the backgroundprocess_uploadCelery task honour. Without the flag the track receives nofederation_object_idand no publication activity.process_uploaddefaults the flag totrueso directory scans keep publishing imported public tracks. - The published
Audioobject carries the track'sdescription(a free-text field settable at upload time or viaPATCH /tracks/{id}) as itscontent: the text is HTML-escaped, http(s) URLs become anchors with scheme-less link text, and#tagsbecomerel="tag"links to this instance's/tags/{name}pages. Tags found in the description are also appended to the object'staglist, and the media download URL is attached as aDocumentwith the audio MIME type. The renderedcontentends with a<p><a href="{track_url}">{artist} - {title}</a></p>link back to the track page (normalize_post_content), andsummaryis left unset: Akkoma and Mastodon both treat a non-emptysummaryas a content warning, so mirroring the post body into it surfaced the raw HTML as a CW header on Akkoma. The object'snameis plain text{artist} - {title}(HTML-escaped): Akkoma already wrapsnamein its own anchor to the objecturl, so an HTMLnameproduced nested anchors. The object also carries a stablepublishedtimestamp (the track'screated_at) — Akkoma falls back to the Unix epoch when it is missing — and theCreateenvelope is stamped with a uniqueidand its ownpublished. Theurllinks also mirrormediaTypeintomimeType— Mastodon's link selection reads the non-standardmimeTypekey and defaults untyped links totext/html, so without it the raw audio download URL would be chosen for display instead of the track page. - Each public lifecycle gets a fresh
Track.federation_object_id(generated when a transition to public is published — album propagation or a trackPATCHwithpublish=true— and re-minted by everyaudiorepublication) so a previousTombstoneat the same URL cannot block re-publication. POST /api/v1/tracks/{id}/publishpublishes a public track to the fediverse at any time (the "Fediverse" tab of the share dialog). Any authenticated user may share a public track as aCreate(Note)post under their own actor; republishing the canonicalCreate(Audio)object (object_type=audio, which re-mintsTrack.federation_object_id) is restricted to the track's owner and admins. The publication activity is attributed to the publisher — itsowner_user_id— who can edit or retract it via the activity endpoints even when they cannot manage the track itself. It accepts an optionalstatus— a one-off post text used as the object'scontentinstead of the storeddescription, run through theprocess_mentionspipeline so@handles become links,Mentiontags, andactivity_mentionsrows — an optionalvisibility(defaultpublic) selecting the post's audience:public/followers/mentionedfederate to their respective audiences, whileprivate/localrecord the activity without delivering it, and an optionalobject_type(defaultnote).noteshares the track as aCreate(Note)— the post body renders on every remote server (Mastodon dropscontentonAudioobjects, which it treats as a converted type) — with the stream embedded as anAudio-typed attachment linked to the track's canonical object when it has one; each share mints its own object id, and the share'surlis that object id rather than the track page: the track URL dereferences to the canonicalAudioobject, so a remote URL lookup (e.g. Mastodon's search box) resolves theAudio, while the share keeps a distinct identity resolvable through its own object URL. The track page still ends the share'scontentas the appended{artist} - {title}link —normalize_post_contenttakes the page URL explicitly (link_href) since it can no longer be read from the object'surl.audiorepublishes the canonicalCreate(Audio)media object. The status is never persisted on the track. Theto/ccaddressing derived from the visibility is applied to both theCreateenvelope and the embedded object. Eachaudiopublication mints a freshfederation_object_idso every post is a distinct remote object.Delete(Tombstone)is sent when a track is made non-public or deleted.- Tracks backed by metadata-only providers (providers declaring
capabilities.limits["federate_audio"] = False, currentlytidalandyoutube) federate without Songhive-hosted audio:track_to_audio_objectandtrack_to_note_objectsuppress the stream/download URL and emit a provider browse link instead (e.g. aLinktotidal.com/browse/track/{id}on theAudioobject, or atext/htmlDocumentattachment onNoteshares), so remote servers never fetch proxied provider bytes. The decision lives infederation/_common.py::provider_browse_linkand is taken at serialization time so local metadata stays complete. - Profile changes (display name, bio, avatar, links) refresh the cached actor
document via
sync_user_actorand are pushed to follower inboxes asUpdate(Person)activities. - Profile links are emitted as
PropertyValueattachments whosevalueis an HTML anchor (<a href="..." rel="me">) so remote servers render them as clickable fields. The URL is only linkified when it is a well-formed http(s) URL with a host and no HTML-breaking characters; malformed values are emitted as escaped plain text instead. Anchor text omits thehttp(s)://scheme for readability. - The actor
summaryis rendered from the bio as escaped HTML with http(s) URLs linkified into anchors (scheme-less link text); sentence punctuation wrapping a URL stays outside the anchor, and URL-looking text that fails validation is left as escaped text. - Following/unfollowing uses standard AP
Follow/Undo(Follow)activities. track.genreis split into multipleHashtagtags on the publishedAudioobject, with spaces converted to underscores. These tags include anhrefpointing at the instance's/tags/{name}page.- Per-actor follower isolation is delegated to pubby's
target_actor_id. - Incoming
Follow/Undo(Follow)activities persist or remove rows in pubby'sfederation_followersstorage; pubby'sInboxProcessoralso retracts the follower when a remote actor'sDeletetargets the sending actor itself — scoped to the actor the processor is bound to. For shared-inbox deliveries that binding is the instance actor, soprocess_incomingextends the retraction: a verified self-Deleteremoves all of the actor's follow records (_retract_shared_inbox_follows). The sweep only runs after pubby's signature verification bound the signer toactivity.actor, and is skipped for blocked/non-allowed domains, which the processor drops before verifying.services/federation.get_actor_followersreturns an actor's followers newest-first byfollowed_atandcount_followers_by_actorprovides per-actor counts, backing thefollowers_countfield onGET /api/v1/users/{username}and the paginatedGET /api/v1/users/{username}/followersJSON endpoint (distinct from the ActivityPubOrderedCollectionserved at/users/{username}/followers). The SPA shows the count on/@{username}and renders follower details at/@{username}/followers. - Follower approval: each user's
followers_approvalpreference (accept,manual, orreject; defaultaccept, editable from/settingsviaPATCH /api/v1/users/me) decides how incomingFollowactivities are handled.process_incomingpasses pubby'sInboxProcessorafollow_policycallback that resolves the target actor URL to the owning user's setting (object follows and the instance actor keep the auto-accept default).acceptstores the follower and repliesAcceptas before;manualstores a pendingFollowRequestin pubby'sfederation_follow_requestsstorage, sends no reply, and flags the resultingfollownotification withfollow_request_pendingso the UI can render accept/reject controls in the notification body;rejectanswers with aRejectactivity and stores nothing (no notification either). Pending requests are owner-only:GET /api/v1/users/me/follow-requestslists them andPOST .../follow-requests/accept|rejectresolve them through pubby'saccept_follow_request/reject_follow_request, which embed the originalFollowin the response, promote approved requests tofederation_followersrows, and hand the reply to thedeliver_activityCelery task for signed delivery. Resolving a request also rewrites the notification payload tofollow_request_status(accepted/rejected) and pushes anotification_updatedevent.Undo(Follow)removes a still-pending request. The SPA exposes requests as an owner-only "Requests" tab on/@{username}/followers; the actor document advertisesmanuallyApprovesFollowers(and the Mastodon-compatible account setslocked) while the policy ismanual. - Object follows: a
Followmay also target a local object rather than an actor — Friendica sendsFollowon a thread's root item (parent-uri) for conversation subscriptions, the FEP-efda "followable objects" pattern. Pubby stores those rows scoped to the object's own URL intarget_actor_id, so they never count as actor followers (the counts above key on the actor URL) andUndo(Follow)retracts them by the same key.Follows targeting remote actors or objects are dropped by the processor without anAccept— the remote server owns their followers collection. Object-scoped rows are read in bulk throughstorage.get_followers_of_targets(...)and their inboxes collected byservices/federation.get_object_follower_inboxes:resolve_audiencefans public activities out to followers of the activity's own object id and everyin_reply_to_activity_idancestor (covering replies to followed threads), and newly materialized public remote replies/quotes are relayed to the same subscribers — the raw activity forwarded verbatim and signed by the nearest local ancestor's owner, or the instance actor when the chain has no local node (federation/incoming._relay_thread_activity). Served object documents advertise theirfollowerscollection (anOrderedCollectionof the subscribers' actor ids) atGET /users/{username}/objects/{object_id}/followers. An object-scopedFollowcreates afollownotification only for the object's owner — a personal-inbox delivery addressed to another user is misaddressed and skipped — withtarget_*payload fields describing the followed object, matching the reply/quote payload convention.
Outbound follows — a local user following a local or remote actor —
are owned by services/follows.py and persisted in the follows table
(models/follow.py), which is the mirroring outbound side of pubby's
inbound follower/request tables. POST /api/v1/users/me/follows accepts a
bare local username, a @user@domain handle, or an actor/profile URL
(remote_content.parse_remote_target splits local from remote inputs,
lookup_remote_actor resolves remote actors through the guarded fetch
and actor cache). Local targets apply the owner's followers_approval
policy immediately — accept writes the pubby follower row and stores an
accepted follow, manual writes a pending FollowRequest and keeps
the row pending, reject fails with 403 — and the target gets the
usual follow notification. Remote targets store a pending row whose
activity_id matches the Follow activity's own id, enqueue it signed
on deliver_activity, and stay pending until the remote server answers:
process_incoming folds inbound Accept/Reject activities back into
the row via apply_follow_decision (matched on the deciding actor being
the follow target plus the wrapped activity id or follower actor URL;
Accept marks the row accepted, Reject drops it). DELETE /api/v1/users/me/follows reverses the relationship: local targets drop
the pubby follower/request rows and retract the notification, remote
targets get a signed Undo(Follow) embedding the original activity
(rebuilt from the stored activity_id) delivered to the recorded inbox.
When a local owner resolves a pending request, the accept/reject
endpoint also updates the local requester's row directly through
apply_local_follow_decision — the federated reply only reaches remote
requesters. GET /api/v1/users/me/follows lists the caller's follows in
all states; GET /api/v1/users/{username}/follows lists accepted
follows publicly and includes pending rows for the owner. Profiles
expose follows_count plus the viewer-relative follow_state, and the
remote-actor endpoint reports follow_state for authenticated callers.
The SPA renders the list at /@{username}/follows (pending badges for
the owner) and follow/unfollow buttons on local and remote profiles.
Inbound remote replies, quotes and announces are materialized into
Activity rows by federation/incoming.py, invoked from
tasks/federation.py's process_incoming for
Create/Update/Delete/Accept/Reject/Announce/Undo
activities after pubby's InboxProcessor has run. Admission is gated on
the outbound-follow graph: a Create is only stored when its
activity.actor is followed by at least one local user
(follows_service.actor_is_followed) — objects fetched explicitly
through remote URL lookup arrive through the dereference path instead,
so nothing from unfollowed actors accumulates locally. Updates to
already-stored objects always apply (the row was admitted by the Create
gate or explicit lookup); an Update for an unknown object — replaying
a missed Create — is subject to the same followed-actor rule, and
Accept/Reject activities additionally answer outbound Follows and
QuoteRequests. A Create whose object replies
to — or quotes — a known activity (matched by source_id or the
/objects/{id} permalink form) becomes a source_type="remote"
reply/quote row attached to the parent's entity, linked through
in_reply_to_activity_id, carrying the raw activity as payload and the
remote actor as source_actor. Quote fields (FEP-0449 quote, Mastodon
quoteUrl, Misskey _misskey_quote) take precedence over inReplyTo,
mirroring pubby's interaction typing. Publicly
addressed objects keep the entity-clamped public visibility; non-public
ones (direct messages, followers-only) are stored with mentioned
visibility when they address at least one local user — through
to/cc/bto/bcc addressees or Mention tags — each resolved local
addressee getting an ActivityMention row so only that audience sees the
post in the listing. Non-public objects addressing no local user are not
materialized. A Create whose object targets no known activity — a
standalone post, or a reply/quote whose parent was never cached — is
instead stored through remote_content.materialize_remote_post: the
object is upserted into remote_objects (same shape as explicit URL
lookup, minus the fetch) and mirrored as an entity_type="remote"
Activity, so a followed actor's outbox shows up on their remote
profile's posts tab. Non-public standalone posts keep private cache
visibility and local activity visibility. Update revises
content, payload, mentions, hashtags and visibility (materializing
posts whose Create was missed, degrading rows that lose their
public audience to mentioned, and refreshing the mirrored
remote_objects row for standalone posts); Delete soft-deletes the
row and tombstones the cache row's unavailable_at, but only for the
recorded source_actor. Attribution is enforced by pubby's attribution.validate:
process_incoming enables strict_attribution on the InboxProcessor,
and materialization/update re-apply pubby.validate_attribution to the
raw activity since the sync runs regardless of the processor's verdict —
the delivering actor must match attributedTo and share the object's
host. An Accept answering a QuoteRequest
we sent (apply_quote_authorization) stamps the issued
QuoteAuthorization id onto the local quote row's stored Create
payload — and fans out an Update — but only when the accepting actor is
the quoted post's author and the instrument/object match the recorded
quote and target. Because the Pubby federation_interactions record
is kept alongside the row, reply/quote counts and listings deduplicate on
object_id; boosts (stored announce rows vs. recorded BOOST
interactions) deduplicate on activity_id — the announce's own id —
since the interaction record has no separate object id. Materialized
replies and quotes render as regular activity
cards and
accept the same interactions as local ones: likes and boosts federate
to the remote author's inbox through fan_out_like_activity/
fan_out_boost_activity, and local replies/quotes address the remote
author via
a Mention tag with inReplyTo/quote fields set to the remote
source_id.
An inbound Announce from a followed actor
(materialize_remote_announce) is stored as a source_type="remote"
announce row attached to the boosted activity's entity through
in_reply_to_activity_id — the same shape boost_activity produces
locally — so the boost surfaces in timelines as an "X boosted" card. A
boosted object unknown locally is first fetched through
remote_content.dereference_remote_object — the guarded remote fetch
that upserts the remote_objects cache row, materializes the object's
mirror Activity, and caches its author actor — and the boosting actor
is refreshed through lookup_remote_actor so the card renders a
profile. Publicly addressed announces inherit the target's visibility;
non-public ones are stored with mentioned visibility only when they
address a local user. Undo(Announce) — whether it wraps the full
announce or only its id — soft-deletes the row, and Delete of an
announce retracts it without tombstoning the boosted object's cache row
(only rows mirroring the deleted object itself tombstone their
remote_objects row). Likes stay interaction-only.
Instance-level actor:
_setup_federation()inapi/app.pyconfigures anApplication-type actor via pubby'sActorConfigand mounts both ActivityPub and Mastodon API compatibility endpoints.
Domain allow/block lists (federation.allowed_instances /
federation.blocked_instances) are enforced at several seams:
POST /users/{username}/inboxandPOST /ap/inboxinapi/routes/federation.pyreject with 403 before queueing (is_domain_allowed). The shared-inbox route is Songhive's own — it shadows pubby's synchronous handler so deliveries queue ontoprocess_incoming(withusername=None) like per-user inboxes: reply materialization and notifications apply there too, whereas pubby's handler stopped at interaction storage and dropped non-publicCreates entirely. For shared-inbox deliveries the notification recipients are resolved from the activity's audience (federation/notifications.resolve_inbox_recipients): local actor URLs into/cc/bto/bcc,Mentiontag targets, stringobjecttargets, and the local owners of objects referenced throughinReplyTo, quote fields, or Like/Announce targets.tasks/federation.py'sprocess_incomingpasses the lists to pubby'sInboxProcessor, which drops the activity before signature verification.- Pubby's outbound fan-out gets the same filtering via
allowed_instances/blocked_instancesonActivityPubHandler(_setup_federationinapi/app.py). tasks/federation.py'sdeliver_activitydrops outbound deliveries to blocked domains before signing.
Remote discovery (explicit lookup):
Songhive supports explicit, bounded lookup of remote actors, activities, and resources — without crawling remote timelines or indexing the fediverse.
services/remote_content.pyclassifies lookup input (@user@domainhandles, actor/profile URLs, object/activity URLs, Songhive-style/tracks/{id}resource URLs) and enforces theremote_search_accesspolicy (disabled/authenticated/public, defaultauthenticated, editable at runtime by admins) plus the federation domain allow/block lists before any network access. Local-domain inputs never hit the network — and skip the policy, since no remote access happens:resolve_local_targetmaps them straight to their SPA route (/{kind}/{id},/@{user},/activities/{id}; object permalinks resolve throughTrack.federation_object_id/Activity.local_object_id).- All remote dereferencing goes through
federation/fetch.py'sguarded_fetch: onlyhttp(s)URLs, DNS answers restricted to globally routable addresses, every redirect target revalidated (max 3 hops), bodies capped at 1 MiB, ActivityPub content types required, 404/410 mapped to "gone". Each outbound request hop is bounded byfederation.fetch_timeout_seconds(default 20s, 1–300;SONGHIVE_FEDERATION__FETCH_TIMEOUT_SECONDS, also admin-editable at runtime). When federation keys exist, each hop is signed with the instance actor key (HTTP signatures are regenerated per hop). - Remote actors resolve through WebFinger and are cached in pubby's
federation_actor_cache— a fetch-on-miss cache with freshness checks and tombstone markers for gone actors. - Remote actor ids are not assumed to carry a username in their path tail:
newer Mastodon versions issue opaque
…/ap/users/{id}URIs. Displayed and routed handles therefore come from the cached actor document'spreferredUsername—services/federation.actor_doc_handlebuilds theuser@domainform, andcached_actor_docs/cached_actor_handlesbatch-resolve it across the actor cache, follower and follow-request tables without any fetch. API responses expose the resolved handle (handle,actor_handle,source_actor_handle) so the SPA never has to parse the actor URL; handle lookup (_find_cached_actor_by_handle) scans the same tables onpreferredUsernameand falls back to matching the actor URL's path tail, so both real handles and legacy numeric-id links resolve to a cached actor before WebFinger is tried. - Remote objects are cached in the
remote_objectstable (canonical URL, activity URL, domain, type, normalized resource kind, denormalized name/summary/content/media URLs, ETag/Last-Modified, content hash,unavailable_attombstone).Create/Announce/Update/Likewrappers are unwrapped,attributedTois validated against the publishing actor, andDelete/Tombstonedocuments mark the cached row unavailable instead of deleting it. - Content objects materialize as
Activityrows withentity_type="remote"pointing at theirRemoteObjectrow andsource_type="remote"/source_id=<canonical URL>; bare resource documents (e.g. a dereferencedAudiotrack) are cached but never become feed entries. Reply/quote parents resolve only against already-cached rows — remote threads are never fetched recursively. - Routes:
GET /api/v1/remote/lookup(explicit lookup → internal URL),GET /api/v1/remote/actors/{handle}(+/activitiesfor cached posts),GET /api/v1/remote/objects/{id}(refreshre-dereferences the canonical URL), andGET /api/v1/remote/{kind}/{id}for cached remote resources. Internal SPA URLs are/@user@domain,/activities/@user@domain/{remote_object_id}, and/remote/{kind}/{id}— remote URLs are never used as navigation targets. - Aggregate search gains a cached-only
remotesection (entities=remoteor theinclude_remoteflag) that never performs network access, plus aremote_availableflag telling the UI whether the caller may run an explicit lookup.
Federated music entities (Funkwhale dialect):
Songhive publishes and consumes a Funkwhale-compatible ActivityPub music dialect, so music entities federate bidirectionally with Funkwhale instances and with other Songhive instances.
federation/serializers.pyserializesArtist,Album,Track,AudioandLibrarydocuments under thehttps://funkwhale.audio/nscontext (MUSIC_ENTITY_CONTEXT): artists and albums carrymusicbrainzId/released/embeddedartists, tracks carryposition/discand a nestedalbum, andAudiouploads carry an integerdurationplusbitrate/size, alibraryreference and the embeddedtrackdocument — the fields Funkwhale's serializers require on import.create_audio_activitypasses the publishing track's library URL (/libraries/{id}for a publicLibrary, else the owner's implicit{actor}/librarycollection) so remote library followers receiveCreate(Audio)deliveries.- Dereferenceable routes in
api/routes/federation.pyserve the music entities as ActivityPub JSON (HTML requests still get the SPA):/artists/{id},/albums/{id}(public albums only),/libraries/{id}for public libraries — the collection index, with?page=Nserving aCollectionPageofAudioitems and/followersserving the object's follower collection — and/users/{username}/library, an implicit followable library of the user's public tracks. Actor documents advertiseendpoints.sharedInboxand alibrarylink to that implicit collection. remote_content.pyrecognizes Songhive resource URLs and Funkwhale's/federation/music/{tracks,albums,artists,libraries,uploads}/{id}and/federation/actors/{name}paths. Music documents classify to a normalizedresource_type(Audio→track,Track→track,Album→album,Artist→artist,Library→library);uploadsURLs normalize totrack. Containment is recorded on the newremote_objects.parent_urlcolumn (Audio→library,Track→album,Album→first artist, collection pages→partOf), and entities embedded in a fetched document (Audio.track,Track.album,Album.artists, pageitems) are cached in the same table — bounded at 120 per fetch — so remote album/library pages render their contents without a fetch per child.Librarydocuments additionally trigger one bounded fetch of their first collection page; a pasted?page=NURL resolves to itspartOfcollection. Funkwhale'saudience-only public addressing is honored (_doc_is_public), and actor URLs that fail to dereference retry once through WebFinger.- Remote resources can be followed object-scoped (FEP-efda style):
POST/DELETE /api/v1/remote/objects/{id}/followstores afollowsrow keyed on the object URL and deliversFollow/Undo(Follow)to the object's controlling actor inbox (actor/attributedTo, e.g. a Funkwhale library's channel actor). InboundAccept/Rejectare matched inapply_follow_decisionagainst the followed object or its owning actor. InboundFollows of a local object delivered to the shared inbox are processed under the object's owning user actor (_object_follow_owner_usernameintasks/federation.py) so theAccept.actormatches the library's controlling actor — which Funkwhale requires. InboundCreate/Announce/Updateadmission accepts activities whose object is — or names as its container (library,context,target) — a followed remote resource, in addition to the followed-actor rule. - Remote resources can be collected:
item_type="remote"onPOST/DELETE /api/v1/collection/{item_type}/{item_id}bookmarks a cachedremote_objectsrow by id — only rows with aresource_typeare collectable, and nothing is copied into local music tables.GET /api/v1/remote/objects?collection=true&resource_type=...lists them (authenticated), andRemoteCollectionSection.vuerenders them inside the Tracks/Albums/Artists/Libraries/Playlists collection views. - Remote media resolution happens at play time, never at cache time:
a remote track is metadata whose media can be a direct
audio/*link (audio_url) or a rendition sibling (Audio/Videodocument embeddingtrack). Rendition rows record that link on the indexedremote_objects.media_of_urlcolumn (Audio→track.id), somedia_of_url == track.canonical_urlresolves without payload scans (a bounded JSON-path fallback covers rows cached before the column existed).GET /api/v1/remote/objects/{id}/streamresolves throughremote_content.resolve_media_url— ownaudio_urlfirst, else a cached rendition — and 302s to the media URL (404 when nothing is playable). Responses advertisestream_urlonly when something is resolvable now; it is computed in one batchedrendition_audio_mapquery per response. This endpoint is the seam for providers whose links expire or require resolution (e.g. future YouTube/Spotify adapters): they plug intoresolve_media_urland clients never handle a remote URL directly. - Remote object responses (
RemoteObjectResponse) carryin_collection,follow_state,parentanditems(cached children), andRemoteResourceView.vuerenders them with follow and collectionEntityActions, a parent link and a contents list.
Dereference stampede defenses (federation/doc_cache.py, built on
pubby.cache) — a boosted post makes hundreds of remote instances
dereference the same actor, object, collection and WebFinger URLs within
seconds. Every AP read endpoint (/users/{u} actor documents and @{u},
objects/{id} and their followers, quote_authorizations,
/activities/{id}, /tracks/{id}, outbox/followers/following
collections, nodeinfo, WebFinger and the music entities) serves through
doc_cache.get_or_render, which delegates to a shared
pubby.cache.DocumentCache: a bounded (10k-entry LRU) TTL cache
(federation.document_cache_ttl_seconds, default 60 s — misses are
capped at 15 s) with single-flight coalescing, so concurrent identical
fetches share one render task (a disconnecting waiter cannot cancel it
for the others), plus stale-if-error serving while an entry is within
ttl * stale_factor. Keys are tuples — ("obj", user, object_id) —
so prefix/segment invalidation is element-wise and cannot collide across
users or endpoints. Renders open their own get_session() session — they
may outlive the request that started them, so they must not borrow the
request-scoped session. Invalidation is generation-guarded: a render in
flight when an invalidation lands cannot re-store a stale result.
Mutating paths call invalidate_actor, invalidate_activity,
invalidate_track, invalidate_keys, invalidate_segment or
invalidate_prefix (profile sync, retractions, cascade deletes,
follow-request approval, local follow/unfollow) with session= so the
drop fires on after_commit — an early invalidation would let a
concurrent render re-cache the pre-commit row; the listener ignores
SAVEPOINT releases (after_commit also fires for nested-transaction
commits) and drains only after the outermost commit. Track mutations
invalidate on the mutation itself, not via publication rows: a
public → private PATCH drops the track's cached page/object/library
documents even when no live create activities exist to retract, and
sync_track_publications drops the object/activity documents whose
stored payloads it rebuilds. Album and library PATCH routes and the
album/artist/library deletion paths drop their own documents the same
way (("album", id), ("artist", id), ("lib", id, …) plus the
visibility-changed tracks). Rare privacy-relevant events that touch
non-user-keyed documents —
admin suspension/unsuspension, payment deactivation, account deletion —
call clear(session=…) instead: ("trackpage", id), ("act", id) and
the lib/artist/album keys are not scoped by username, so
/tracks/{id} and /activities/{id} would otherwise keep serving the
user's documents for up to the TTL. The TTL bounds whatever cross-process
staleness remains (best-effort). Responses carry Cache-Control: public, max-age=<remaining ttl> — the cache entry's remaining
freshness, not a restarted full TTL, so downstream caches cannot extend
an aged document's lifetime; an exhausted-freshness response carries
max-age=0, must-revalidate and a disabled cache (ttl = 0) carries
no-store — plus Vary: Accept and an ETag
(conditional requests answer 304), so shared caches and the nginx edge
cache can absorb repeats before they reach the app. The nginx layer
(docker/nginx.conf) proxy_caches only anonymous AP-JSON requests —
HTML variants and credential-bearing requests bypass it — mirrors the
app's Accept negotiation into three key classes (*/* counts as
HTML-capable, matching _accepts_html), ignores upstream Vary (the
normalized class is already in the key; honoring it would fragment on
every remote's raw Accept string), and single-flights misses with
proxy_cache_lock; the optional per-IP limit_req (429) requires
set_real_ip_from for the front proxy first and never applies to inbox
POSTs. The edge cache sits outside every invalidation hook — app-side
clears do not reach it — so every cacheable response also carries
X-Accel-Expires bounded by both the remaining in-process freshness and
federation.edge_cache_ttl_seconds (default 15 s, matching the
proxy_cache_valid fallback): the edge TTL can never outlive the app's
freshness policy, and stale serving is disabled at the edge so the
staleness window after a suspension or deletion is hard-bounded at that
cap. Underneath the cache, init_db(url) builds a
bounded AsyncAdaptedQueuePool engine per event loop (asyncpg
connections are loop-bound; the process runs several) honoring
database.pool_size / max_overflow / pool_timeout / pool_recycle /
pool_pre_ping — pool-timeout waits and Postgres-side
too_many_connections both fail fast with 503 + Retry-After instead
of hanging. Short-lived Celery/CLI task loops cap their async pool at
1+2 via database_task_engine_kwargs. The same DocumentCache instance
is passed to bind_activitypub/ActivityPubHandler, so pubby's own
adapter routes (/ap/*, instance actor, nodeinfo) share it and its
mutation hooks invalidate it — pubby's route keys live under a
("pubby", …) namespace so neither side's invalidations touch the
other's entries (Songhive's WebFinger route builds its own
("webfinger", …) key from pubby's normalize_webfinger_resource
rather than pubby's namespaced webfinger_key).
federation/storage.py's get_federation_storage
delegates to pubby's memoized get_db_storage — the same sync engine is
reused process-wide instead of building a new engine and running
create_all per call; its pool is fixed at pool_size=2, max_overflow=3 (pubby calls are short), which also bounds each Celery
prefork child.
Content Moderation
Reportmodel stores user-submitted content flags (target type/id, reason, description, status, reviewer). Actor reports (target_typeuser/actor) accept a username,@user@domainhandle or actor URL, are normalized totarget_type="user", and store the reported account's canonical actor URL intarget_actor_url(target_idholds the local user id or the remote actor URL; legacy callers may still pass a bare user id). Theforwardedflag records whether the report was relayed to a remote instance.- Public submission:
POST /api/v1/reports.forward: trueon a remote actor report enqueues an ActivityPubFlagactivity to the reported actor's inbox viatasks.federation.deliver_activity(requires the reporter's federation keys; otherwise the report is stored withforwarded=false). - Every report notifies all active admins (except the reporter) through a
reportnotification carryingreport_id/target_type/target_actor_url/reason; the frontend links it to/admin/reports. - Admin review:
GET/PUT /api/v1/admin/reports. Responses includereporter_username,target_actor_urlandforwarded. AuditLogrecords administrative and security-relevant actions (actor, target type/id, IP address, JSON details).target_typevalues are defined by theAuditTargetTypeenum (models/audit_log.py) —log_actiononly accepts enum members, andGET /api/v1/admin/audit/target-typesexposes them to the admin UI's target-type filter.
Actor and instance moderation
Mastodon-style moderation lives in three tables (models/moderation.py) and
one service (services/moderation.py):
UserModeration— per-usermute/blockon any actor (local or remote), keyed on the canonical actor URL (user.actor_url, or theurn:songhive:user:<name>fallback when federation is disabled). A mute is one-way: the target's activities leave the muter's feeds but the muter's activities still reach the target. A block is reciprocal: it hides both directions, prevents interaction both ways, and severs every follow relationship between the pair (sever_block_relationshipremoves thefollowsrows and the pubby follower/request records).AdminUserModeration— instance-widelimit/suspendon an actor with an optionalreason, unique per actor URL (applying a new action upgrades the existing row).limitforces the actor's follows of local users through manual approval (follow_userpassesforce_manual) and gates their activities to their own followers — on timelines and on their profile page, which shows a "limited by the moderators" notice and a "show anyway" button that re-fetches the timeline withreveal=1(therevealquery param lifts only the followers-only gate for that actor; hidden/defederated content stays excluded).suspendadditionally blocks all interaction (assert_interaction_allowed), drops inbound activities (tasks/federation.process_incoming), removes the actor from notification fan-out, severs every local follow relationship (sever_relationships), and hides all of the actor's content.InstanceModeration— per-domaindefederate/followers_onlypolicies with an optionalreason, layered over the configuredfederation.allowed_instances/blocked_instanceslists. Because the synchronous federation paths (Celery delivery, pubby inbox resolution, remote-actor search) cannot query the database,load_instance_policiesinstalls a process-local snapshot (federation.set_db_instance_policies) thatis_domain_blocked/db_domain_policyconsult — adefederaterow behaves exactly like a configured block, whilefollowers_onlygates the domain's activities to local users who follow the author.
Visibility is enforced through moderation_context (one small set of queries
per request) + moderation_filter/activity_hidden, applied to activity
lists, single-activity views, timelines, search results, and remote-actor
lookups. Interaction endpoints run assert_not_suspended/
assert_interaction_allowed; notification fan-out checks
notification_suppressed, which also drops filterable notifications
(everything but report) from limited actors and followers-only-domain
actors when the recipient does not follow them — Mastodon's
for_limited_accounts: drop policy. Federated dereference endpoints return 404 for
suspended owners, and defederated domains are dropped inbound (inbox
processing, remote materialization) and outbound (delivery domain checks,
audience resolution).
API surface:
GET/POST/DELETE /api/v1/users/me/mutesand.../blocks— personal moderation lists (ModeratedActorResponsecarries a display snapshot).GET/POST/DELETE /api/v1/admin/moderation/users— admin limit/suspend.GET/POST/DELETE /api/v1/admin/moderation/instances— admin domain policies.GET /api/v1/instance/domain_blocks— Mastodon-compatible transparency endpoint mappingdefederate→suspendandfollowers_only→silence.GET /api/v1/users/{username}and/api/v1/remote/actors/{handle}reportmuted/blocked(viewer-relative) andlimited/suspended(instance) flags; profile pages show the badges plus mute/block actions for signed-in viewers and limit/suspend/clear actions for admins./settings?tab=moderationlists the viewer's mutes/blocks with undo;/admin/moderationmanages all admin actions with optional reasons.- All admin mutations record
moderation.<verb>AuditLogentries.
Sharing
Two independent mechanisms for sharing private content:
| Mechanism | Model | Route prefix | Description |
|---|---|---|---|
| Share grant | ShareGrant |
/api/v1/shares |
Owner grants a named user access |
| Share URL | ShareToken |
/api/v1/share-urls |
Revocable short link; token hash stored |
| URL resolver | — | /api/v1/share/{token} |
Resolves raw token → HTML, JSON, or direct audio download |
The URL resolver is content-negotiated:
- Browsers and crawlers receive a rendered HTML preview page with OpenGraph metadata, an audio player for tracks, and track listings for albums, playlists, libraries, and artists.
- API clients that send
Accept: application/json(including the web UI's share preview) still receive a302redirect to the item's public JSON endpoint, with a short-livedshare_tokencookie. - Audio file shares redirect directly to the file download URL, and
?download=trueforces a direct audio download for tracks or files.
The ACL service (services/acl.py) checks grants and tokens transparently
via the require_access FastAPI dependency used by resource routes.
Embeddables
Public entities can also be embedded into third-party pages from the share
dialog's "Embed" tab (components/share/EmbedPanel.vue). Four snippet
formats are generated client-side (utils/embed.ts):
- Tracks default to a standalone
<audio>tag wrapped in a<p>with the track metadata linked back to the Songhive page. For albums, artists, playlists, and libraries this is replaced by a<div>listing one<audio>element per track — only offered when the collection holds fewer than 250 elements, and limited to public tracks with a playableaudio_url. - A Markdown link
[{artist} - {title}]({url})(title or filename fallbacks for tracks, the entity name for collections). - A
<script>embed: adata-songhive-embedplaceholder rendered byfrontend/public/embed.js(served as/embed.js) into an iframe of the SPA's standalone/embed/{type}/{id}route (views/EmbedView.vue) — a compact track player for tracks and an expandable tracklist for collections. The embed page reports its content height viapostMessageand the script resizes the iframe to match. - A no-JS
<iframe>fallback pointing at the same/embed/{type}/{id}route with fixed heights.
Embed URLs load anonymously, so the tab only offers snippets for public items (owners see a publish-first hint otherwise).
Feeds
api/routes/feeds.py serves RSS 2.0 and Atom 1.0 documents under /feeds
(outside /api/v1, like /webmentions and /ap), rendered by
services/feeds.py with standard-library XML escaping — no feed library is
involved. Each feed is available in both formats by swapping the URL suffix:
| Feed URL | Content |
|---|---|
/feeds/users/{username}.{fmt} |
The user's latest posts (mode="posts" timeline) |
/feeds/{plural}/{id}/activities.{fmt} |
Activities attached to a track, album, artist, playlist, library or radio |
/feeds/artists/{id}.{fmt} |
Latest releases: albums plus standalone tracks merged by date |
/feeds/playlists/{id}.{fmt} |
Tracks most recently added to the playlist |
/feeds/libraries/{id}.{fmt} |
Tracks most recently added to the library |
/feeds/tags/{name}.{fmt} |
Newest entities and activities carrying the tag |
/feeds/genres/{name}.{fmt} |
Newest tracks and albums in the genre |
{fmt} is rss or atom (served as application/rss+xml /
application/atom+xml); ?limit= may reduce the page size but never
exceed feeds.max_items (default 20, range 1–500). The whole subsystem is
gated on feeds.enabled. Feed queries reuse the same ACL predicates as the
JSON API, so anonymous readers only see public content and authenticated
requests also cover local visibility and share grants; a valid share
token still grants access to a private collection's feed. For the artist
releases feed, a track that belongs to an album visible to the requester is
not listed separately — the album entry already represents it. Track items
carry an <enclosure>/Atom rel="enclosure" link pointing at the public
/api/v1/stream/{id} URL when the track has audio.
Feed discovery happens server-side so non-browser clients can find the
URLs: api/semantic_meta.py appends
<link rel="alternate" type="application/rss+xml|atom+xml"> tags to the
semantic <head> tags it injects into the SPA shell for object pages (and
profile_pages does the same for /@{username}), mirroring whatever feed
the page exposes — track and album pages advertise their activities feed,
/activities sub-pages of artists/playlists/libraries switch to the
activities variant. On the frontend, utils/feeds.ts builds the same URLs,
composables/useFeedLinks.ts keeps document.head's feed alternates in
sync after client-side navigation (replacing the tags the server injected
for the landing page and removing them on feed-less pages), and
components/ui/FeedButton.vue renders an RSS/Atom menu on every
feed-backed page (entity headers, activity pages, tag/genre detail and
user profiles).
Notifications
models/notification.py defines Notification (recipient user_id,
type, optional actor_url/source_url, JSON payload,
delivered_targets, seen_at, digest_sent_at) and
NotificationPreference (unique per (user_id, type) with in_app,
email, and email_digest toggles). The notification types are
follow, like, boost, quote, reply, mention, share,
webmention, activity, report (admin-only: a user report was
filed — links to /admin/reports), and download (a bulk-download ZIP
archive finished building — the payload carries archive_id, label,
size, and item_errors, and source_url points at /downloads).
ActivitySubscription rows
record that a local user wants an activity notification for every
activity another actor authors — the "bell" toggle on a user profile.
Local users are referenced through target_user_id, remote actors through
target_actor_url (exactly one is set per row, each unique per
(user_id, target) pair, no self-subscriptions).
services/notifications.py creates notifications: it resolves the
recipient's per-type targets (defaults when no preference row exists are
in-app on, email and digest off), skips creation entirely when all targets
are disabled, and deduplicates on (user, type, actor_url, source_url)
against existing unseen rows. When in_app is enabled it pushes a
notification event over the WebSocket bus via
EventWebSocket.send_to_user, delivered only to the recipient's
connections. When email is enabled and the recipient has a verified email
address it enqueues tasks/email.py's send_notification_email. Rows with
email_digest enabled are collected by the scheduled
send_notification_digest task, which groups pending rows per user, sends
one plain-text email, and stamps digest_sent_at only after a successful
send. purge_seen_notifications deletes seen rows older than
notifications.retention_days (default 90). Hook and email failures are
logged and swallowed so they never break the originating operation.
The same module also removes and rewrites notifications:
delete_notifications and clear_notifications handle recipient-scoped
manual deletion, while retract_notifications (filtering on recipient,
type, actor URL variants, and source_url) and
retract_notifications_referencing (matching source_url,
payload.target_url, and payload.activity_id; dialect-aware for
PostgreSQL vs SQLite/MySQL JSON extraction) clean up rows whose
underlying event was undone or deleted. Every deletion path pushes a
notification_deleted event ({"ids": [...]}) to each affected
recipient over the WebSocket bus. update_notifications_referencing
rewrites the managed snapshot fields of rows referencing a set of
activity/object URLs, update_notifications_from_actor refreshes actor
metadata (actor_name/actor_display_name/actor_avatar_url) on rows
produced by a given actor, and refresh_notifications_for_item updates
item_title/target_item_title/track_title on rows referencing a
local item; each pushes
a notification_updated event carrying the changed rows. Updates only
touch managed fields — unrelated payload keys and read/unread state are
preserved, and missing source fields remove stale keys. Service helpers
flush but never commit.
REST API (api/routes/notifications.py; every endpoint is authenticated
and scoped to the current user):
| Route | Description |
|---|---|
GET /api/v1/notifications/ |
Newest-first list, seen filter, type CSV allowlist, limit/offset + X-Total-Count |
GET /api/v1/notifications/unread-count |
Unseen count for the nav badge |
POST /api/v1/notifications/seen |
Mark ids seen |
POST /api/v1/notifications/unseen |
Mark ids unseen |
POST /api/v1/notifications/seen-all |
Mark all current-user rows seen |
DELETE /api/v1/notifications/{id} |
Delete one owned notification |
POST /api/v1/notifications/delete |
Bulk delete owned ids (max 500) |
POST /api/v1/notifications/clear |
Delete all current-user rows |
GET/PUT /api/v1/notifications/preferences |
Merged per-type delivery matrix |
POST /api/v1/admin/notifications/purge performs a manual retention purge
and records a notification.purge audit row with the deleted count and
retention window; songhive admin purge-notifications runs the same service
from the CLI, and the /admin/tasks page exposes the endpoint as a
confirmation-protected action.
Creation hooks: favoriting another user's track
(api/routes/favorites.py) creates a like, and liking another local
user's activity (services/activities.py's like_activity, reached via
POST /activities/{id}/like) notifies its owner_user_id; creating a new
share grant (api/routes/shares.py via services/sharing.py's
(grant, created) return) creates a share. The federation inbox path
(tasks/federation.py → federation/notifications.py) maps Follow/Like/
Announce/Create/QuoteRequest activities for each resolved local recipient
— the addressed user for per-user inboxes, resolve_inbox_recipients'
audience resolution for shared-inbox deliveries — quote wins over
reply when a Create is both, only the target's owner gets the
reply/quote notification (a reply or quote of someone else's post
that merely tags the recipient stays a mention, which is otherwise
suppressed once a reply/quote notification for the same note fired),
and a Like/Announce likewise only notifies the target's owner (a
followed actor's like or boost of somebody else's — or a remote — post
notifies nobody: the interaction is pubby's federation_interactions
record plus, for followed actors' announces, a materialized announce
row) — and stamps matching ActivityMention.notified_at rows. A FEP-044f
QuoteRequest (auto-approved by pubby, which stores a dereferenceable
QuoteAuthorization and answers Accept) also yields a quote
notification when its object resolves to a local post owned by the
recipient — the quote's own Create may never reach the inbox — with the
instrument id recorded as source_url. Because a quote can arrive
twice (request then Create, or the same Create on two inboxes) and
the first row may already be seen, quote notifications deduplicate and
merge payloads on (recipient, actor, quoting object) regardless of
seen state, preserving the existing row and its read marker
(_create_or_update_quote_notification).
Activity subscriptions fan out through
services/activities.py's _notify_activity_subscribers, invoked by every
local producer path (create_local_activity for create types,
create_status, record_track_publication, like_activity,
boost_activity, reply_to_activity, quote_activity). Each subscriber
is filtered through can_view_activity — mentioned/private posts and
inaccessible entities never leak — and the author, the interaction
target's owner, and already-mentioned users are skipped so the activity
row stays a fallback rather than a second notification for an event a
specific reply/quote/like/boost/mention row already covers.
The payload snapshots the authored activity (or, for authored
likes/boosts, the reacted one) with the same object_*/target_*/
item_* fields the other hooks produce, plus activity_type so clients
can phrase the action ("shared a post", "liked a track").
Remote actors fan out the same way through
notify_remote_activity_subscribers, keyed on the materialized row's
source_actor: it is invoked by the inbox materializers —
_materialize_remote_object (replies and quotes),
materialize_remote_announce, and materialize_remote_post via a
notify_subscribers flag on _materialize_remote_activity, so only
inbox-delivered activities notify while explicit remote-object lookups do
not. Actor display fields come from the federation actor cache
(resolve_source_actor_profiles) with a user@domain handle fallback,
and remote-entity rows link the local /activities/{id} page since they
resolve no local entity. A remote actor Delete drops the subscriptions
targeting them.
Subscription state is toggled via
POST/DELETE /api/v1/users/{username}/activity-subscription for local
users and POST/DELETE /api/v1/remote/actors/{handle}/activity-subscription
for remote actors; subscribing also follows the target on a best-effort
basis (a failed follow does not fail the subscription) and the response
reports the resulting follow_state. The viewer's state is surfaced as
activity_subscribed on the public profile and remote-actor responses;
account deletion drops the rows in both directions, including any keyed on
the deleted user's actor URL.
Notification payloads are denormalized at creation so rows stay renderable
after the source object disappears. Every hook records the actor's
actor_name/actor_display_name/actor_avatar_url (the federated values
come from pubby's cached actor document); likes, boosts, and shares record
item_type/item_id/item_title/local_url when the object resolves to a
local entity (federation/notifications.py's _resolve_local_object
matches {actor}/objects/{id} against Track.federation_object_id and
Activity.local_object_id/source_id, and /{plural}/{id} URLs on the
instance domain). Objects resolving to a local Activity additionally
record object_activity_id, the ActivityStreams object_type (Note,
Audio, …) and object_page_url — the activity's own page, i.e. the
entity's activity feed or the author's profile for user entities
(services/activities.py's activity_page_url, shared by
_entity_link_fields and the object-permalink browser redirects). user
entities have no item page, so item_type/item_id are omitted there and
local_url points at the author's /@username profile. Create payloads carry a _note_snapshot —
object_content (capped raw HTML), object_summary, object_name,
object_url, published, object_mentions — and replies/quotes add
target_url plus target_* fields for a resolved local target. The
RE: <link> quote fallback remote servers embed for non-quote-aware
clients (a quote-inline element for Mastodon/Akkoma, a bare RE: <url>
tail for Misskey/Threads) is stripped from stored content and snapshots
(strip_quote_fallback) — the quoted activity renders through the
in_reply_to_activity_id embed instead.
Retraction mirrors creation so notifications don't outlive their event:
unfavoriting a track retracts the owner's like notification, retracting a
like activity removes the notification it produced (matched via
payload.activity_id), revoking a share grant retracts the grantee's
share, retract_activity and cascade_delete_entity retract
notifications referencing the removed activity/entity, and deleting a user
retracts the notifications they produced elsewhere. On the inbox path, federation/notifications.py's
retract_inbox_notifications handles Undo (Follow/Like/Announce,
including undo-by-activity-id via payload.activity_id) and Delete
(object or actor) scoped to the addressed recipient.
Edits propagate too: update_activity and sync_track_publications
refresh the object snapshot on notifications referencing the rebuilt
object (retracting rows whose mention/reply/quote basis no longer
holds), and track renames refresh the stored item titles. On the
inbox path update_inbox_notifications handles Update activities —
a Note update refreshes the snapshot when the notification is still
applicable and retracts it when the recipient is no longer mentioned or
the reply/quote target changed away, while an actor Update refreshes
the stored actor metadata.
The frontend stores/notifications.ts owns the unread count, the paginated
list, and optimistic seen/unseen and deletion updates (with rollback +
toast on error). The list composes an all/unread read-state filter with a
multi-select type allowlist (empty = all types) sent as the type CSV param;
incoming events of filtered-out types still bump the badge but stay off the
list.
It registers notification, notification_deleted, and
notification_updated handlers on the
shared eventBus singleton
(api/ws.ts), connected when the auth store reports authenticated and
disconnected on logout; each received notification also raises an info
toast naming the actor and action, while deletions prune matching rows
and updates replace changed rows in place, both silently. AppLayout renders a 99+-capped badge on the
/notifications nav item; views/NotificationsView.vue auto-marks visible
rows seen via a debounced IntersectionObserver batch that defers to rows
the user just toggled back to unseen, and offers per-row dismissal plus a
TrackList-style selection mode for bulk delete/read/unread and a
confirmation-protected "Clear all". Each row's header links the actor name
to the actor's internal profile (local or remote /@name@host) and the
action text separately
to the referenced object — for likes/boosts the object_page_url of the
reacted activity. The action text names the interacted entity
(utils/notifications.ts, shared with the WebSocket toast): likes/boosts
on Audio objects or resolved item_types render as "liked/boosted your
track/album/…", replies/quotes do the same through their target_*
fields, and shares name the granted item ("shared an album with you"),
while Note objects and unresolved remote objects keep the
generic "post" wording. Rows render context cards from the
payload: follows (and unresolved likes/boosts) show
components/notifications/NotificationActorCard.vue (all actors route to
the internal profile — /@name locally, /@name@host for remote actors,
whose profile keeps the link out to the origin site),
mentions/replies/quotes embed a
read-only ActivityCard fed by the note snapshot (remote HTML reduced to
plain text), likes/boosts on Note objects fetch and embed the real
ActivityCard through NotificationActivityCard.vue (GET /api/v1/activities/{id}), and shares plus likes/boosts on Audio objects
(or failed activity fetches) show
NotificationItemCard.vue, which fills in title and cover art through the
deduplicating composables/useItemSummary.ts cache when the payload lacks
them. The /settings?tab=notifications
profile tab (views/profile/NotificationSettings.vue) edits the per-type
in-app/email/digest matrix; the email columns are disabled until the
account email is verified.
Mentions archive
Mention notifications are dismissible and purged by retention, and they
are never created at all when the recipient's delivery preferences drop
in-app rows or a reply/quote notification already covered the note — so
they cannot answer "which activities ever mentioned me". The mention
archive closes that gap: models/mention_record.py defines
MentionRecord (recipient user_id, source, source_url,
activity_id, actor_url, visibility, JSON payload), a permanent
per-user record of every activity that addressed them, backed by
services/mention_records.py. Rows are keyed on
(user_id, source, source_url) so re-delivery and edits upsert instead
of duplicating, and payload snapshots the same render fields the
matching notification carries (actor identity, object_* snapshot,
target_*/item_*/local_url link fields) so records stay renderable
after the source object disappears.
source identifies the pipeline that produced the mention and each one
writes records independently of notification delivery:
local— activities authored on this instance (services/activities.py's_record_activity_mentions, run bycreate_local_activity,create_status,reply_to_activity, andquote_activity). Every mentioned local user is recorded — including self-mentions and recipients already covered by a reply/quote notification — with the activity'ssource_idassource_url, its materializedactivity_id, and its storedvisibility.activitypub— objects received through the federated inbox (federation/notifications.py's_record_inbox_mention, called fromcreate_inbox_notificationswhenever the note'sMentiontags address the recipient). Visibility is classified from the object's addressing:publicwhen it addresses the public collection,followerswhen the author's followers collection is addressed,mentionedotherwise.webmention— incoming Webmentions materialized into activities (webmentions/service.py's_record_webmention, called frommaterialize_webmentionfor the target's owner). The deterministicurn:songhive:webmention:<hash>activitysource_idkeys the record, so re-sent mentions refresh it and the target's visibility is mirrored.
The lifecycle mirrors the notification one so records don't outlive
their event: local edits (update_activity) re-resolve mentions and
delete records for recipients the edit dropped, retract_activity and
cascade_visibility_update remove or reclassify records with the
activity, federated Undo/Delete retractions
(retract_inbox_notifications) remove records referencing the undone or
deleted object/activity — or every record a deleted actor produced —
and Update handling (update_inbox_notifications) merges the revised
snapshot into surviving records, drops records whose Mention tag was
edited out, and refreshes actor_* fields when the actor document
itself is updated. Webmention retractions flow through
retract_activity. Deleting a user removes their archive and the
records their actor URL produced (users/manager.py's
_remove_user_references).
REST API (api/routes/mentions.py): GET /api/v1/mentions/ is
authenticated and scoped to the current user, newest-first, with
limit/offset + X-Total-Count pagination, a source CSV allowlist
(local, activitypub, webmention; unknown values ignored), and
visibility=private restricting to non-public records.
The frontend /mentions route (views/MentionsView.vue, authenticated)
lists the archive with source and all/private filter button groups and
load-more pagination. Each row reuses the notification card pipeline —
a real ActivityCard through NotificationActivityCard when
activity_id/object_activity_id resolves to a stored activity, a
read-only snapshot ActivityCard for unmaterialized remote notes, then
NotificationItemCard/NotificationActorCard fallbacks — and badges
the source, the Webmention type, and a lock for non-public records.
Task Queue (Celery)
All background work is handled by Celery workers. Redis is the broker (db 1) and result backend (db 2).
| Task module | Responsibilities |
|---|---|
tasks/import_.py |
File processing, tag extraction, track/album/artist upsert |
tasks/federation.py |
Activity delivery, inbox processing, key provisioning, remote-activity pruning |
tasks/transcoding.py |
Pre-transcode to common formats, cache result |
tasks/email.py |
Verification, password-reset, notification emails |
tasks/notifications.py |
Notification digest + seen-notification retention purge (scheduled) |
tasks/musicbrainz.py |
MusicBrainz + Cover Art Archive metadata enrichment |
tasks/images.py |
Artist image + Cover Art Archive cover enrichment |
tasks/preview_cards.py |
Fetch + cache per-URL link-preview cards for activities |
tasks/downloads.py |
ZIP archive builds behind a Redis concurrency semaphore, archive retention cleanup (hourly) |
tasks/storage.py |
Orphaned StoredFile GC, audio-only hash rehash (scheduled) |
The cleanup_orphaned_files_schedule config accepts any 5-field cron
expression. The notification digest and retention purge run on fixed daily
crontabs at notifications.digest_hour (default 8 AM) and
notifications.purge_hour (default 3 AM).
songhive.tasks.federation.prune_remote_activities removes stale remote
content through services.remote_content.prune_stale_remote_activities:
source_type="remote" activities whose published_at is older than
federation.remote_activity_retention_days (default 30, overridable via
the task's older_than_days argument) and that no local user has
interacted with — a live like/boost/reply/quote child reached through
in_reply_to_activity_id protects the whole thread, so the eligible set
is computed to a fixpoint. Mirrored remote_objects cache rows are
removed with their activities and can always be re-fetched via remote
URL search; bare remote resources without an activity row are never
touched. The task runs manually via POST /api/v1/admin/federation/prune-remote-activities (audited as
federation.prune_remote_activities), songhive admin prune-remote-activities, or the admin Tasks page — all with dry-run
support — and only joins the beat schedule when
federation.remote_activity_prune_schedule holds a 5-field cron
expression.
Each task's async work is executed with asyncio.run(...). Because
asyncpg connections are bound to the event loop that created them, every
task disposes the shared async engine and resets the global session factory
before the loop closes, ensuring the next task gets a fresh pool.
SMTP-based email is configured via the email config section. Celery tasks
in tasks/email.py enqueue verification, password-reset, and individual
notification messages asynchronously; tasks/notifications.py sends the
per-user daily digest. EmailNotConfiguredError is raised when the SMTP
host or from_address is missing.
Track Metadata and Audio-Only Content Hashing
Uploaded audio files are deduplicated and addressed by an audio-only SHA-256
hash. The hash is computed with ffmpeg -map 0:a -c copy -f streamhash over
the raw audio bitstream, ignoring container metadata, embedded tags, and cover
art. This means:
- Two files containing the same recording but with different tags have the same
hash and share a single
StoredFilerow and storage path. - Tags and cover art can be rewritten in place without changing
sha256orstorage_path; onlyStoredFile.sizeis updated.
The tag rewrite is performed by the sync_track_tags Celery task
(songhive/tasks/tags.py). The task:
- Acquires a Redis lock at
sync_tags:{track_id}(nx=True,ex=300). - Loads the track with its artist, album, audio file, track image, and album cover relations.
- Resolves cover art in this order:
- Track
image_file_id - Album
cover_file_id - No cover
- Track
- Retrieves the audio file locally and writes the current DB metadata into the
embedded tags using
mutagen. - Reconciles the track's
Genreassociations fromtrack.genreor from the parent album when the track has no explicit genre, creates the corresponding tag associations viagenres_to_tags, and propagates the album's genre from the intersection of its tracks' explicit genres. - Updates
StoredFile.size. For S3, re-uploads the rewritten file to the same key; for local storage, the file is already in place. - Releases the lock.
Tag sync is triggered automatically by metadata-mutating API operations
(track/album/artist PATCH, track/album cover upload and delete) and by
successful MusicBrainz enrichment. Manual bulk triggers are provided by
POST /api/v1/admin/sync-tags and songhive admin sync-tags (with optional
--track-id, --album-id, --artist-id, --library-id, or --all); the
admin web UI exposes the same options under /admin/tasks.
Image enrichment (artist images + album covers) is normally triggered
automatically by MusicBrainz metadata enrichment. It can also be run manually
in bulk via POST /api/v1/admin/enrich-images and
songhive admin enrich-images (with --artist-id, --album-id, or --all,
plus --force to re-process already-enriched entities and --dry-run to
preview counts).
Artist images are resolved in this order: (1) MusicBrainz image URL
relationships, (2) Wikidata P18 image claims via the artist's Wikidata
relationship, (3) archived/known image hosts such as web.archive.org
(Spotify CDN) and i.scdn.co. If no image can be downloaded, the artist is
not marked as enriched, so future syncs will continue to retry.
To migrate an existing library that was stored before audio-only hashing, run
songhive admin rehash-audio (with --dry-run to preview) or use
POST /api/v1/admin/rehash-audio from the admin UI. The task re-hashes audio
files, moves/renames the backing files to the new hash-based paths, and merges
duplicate StoredFile rows.
Federation keys and actor URLs can be back-filled with
songhive admin provision-federation-keys or POST /api/v1/admin/provision-federation-keys.
Admins can also inspect and control the Celery worker pool from the UI. The
GET /api/v1/admin/celery/tasks endpoint lists all tasks currently running on
workers (including per-task runtime, worker, args and kwargs), while
POST /api/v1/admin/celery/terminate accepts a list of task ids and revokes them
with terminate=True so they are killed on the worker(s) that are running them.
The admin UI exposes these under /admin/celery with bulk selection.
Storage Backends
| Backend | Class | Description |
|---|---|---|
| Local | LocalStorage |
Files stored under storage.local_path |
| S3 | S3Storage |
S3-compatible object storage; optional CDN prefix |
Backend is selected via storage.backend config. Both implement the abstract
StorageBackend interface (storage/base.py). FileSizeLimitExceededError
is raised when an upload exceeds storage.max_upload_size.
WebSockets
ws/events.py implements a Tornado WebSocketHandler that:
- Validates the
Originheader againstconfig.server.cors_origins; origins matching the requestHostare always allowed since the SPA and/wsare normally served same-origin (nginx proxies both, and Vite forwards/wsin development). - Authenticates the connecting user via a JWT access token: the
?token=query parameter (API clients) takes precedence, then theaccess_tokenHttpOnlycookie, which browsers send automatically on same-origin handshakes. - Broadcasts real-time events (import progress, federation notifications, etc.) to authenticated clients.
EventWebSocket.send_to_user(user_id, event_type, data)serializes events the same way asbroadcastbut delivers them only to connections owned by the given user; it is used for the targetednotificationevent pushed by the notification service.
broadcast/send_to_user can only see the connections living in their own
process, so both also publish an envelope to the Redis pub/sub channel
songhive:ws-events (via a synchronous client — Celery tasks run each job in
a fresh asyncio.run loop). The Tornado process runs ws_event_subscriber
on its IOLoop, which delivers envelopes from other processes through
EventWebSocket.deliver_envelope; envelopes stamped with this process's id
are skipped since they were already delivered locally. This is what makes
notifications created in the Celery worker (federated follows, likes, etc.)
reach connected clients live.
The frontend EventBus (frontend/src/api/ws.ts) reconnects dropped sockets
with exponential backoff (1s up to 30s). Since the server authenticates after
the upgrade, the backoff only resets once a connection has stayed open for a
while; a close with code 4001 (unauthenticated) additionally triggers an
access-token refresh so the retry uses fresh credentials instead of looping
on a stale token.
Because authentication happens after the handshake, the browser sees an
accepted socket that is then closed — which resets the reconnect backoff of
older clients on every attempt. To bound that churn server-side, the handler
tracks authentication failures per (remote IP, token digest): the first
failure still closes immediately with 4001, but each consecutive failure
delays the close exponentially (1s, 2s, …, capped at 30s).
Frontend
Vue.js 3 + TypeScript SPA, bundled with Vite.
| File/Dir | Role |
|---|---|
frontend/src/main.ts |
App bootstrap, Pinia + i18n + router mount, theme apply |
frontend/src/App.vue |
Root component (<RouterView />) |
frontend/src/router/ |
Vue Router (history mode) with global auth/admin guard |
frontend/src/stores/ |
Pinia stores (auth, theme, toast, confirm, player, playback session, outputs) |
frontend/src/components/ui/ |
Headless base components (button, input, select, avatar, table, pagination, search, context menu, entity actions) |
frontend/src/components/feedback/ |
Toast, banner, spinner, skeleton, modal, confirm dialog |
frontend/src/components/entity/ |
Reusable entity grid/list components (e.g. BulkEditableGrid for bulk selection and deletion) |
frontend/src/components/activities/ |
Activity feed components (ActivityFeed filter tabs + cursor pagination, ActivityCard, ActivityEditModal) backed by stores/activities.ts and api/activities.ts. ActivityAudioPlayer renders Audio attachments as a styled inline player (artwork — attachment image, then the author's avatar, then a note icon — title/artist/album, seek and volume controls) instead of the browser-default element; its "Play in player"/"Add to queue" actions resolve local songhive:trackId attachments through the tracks API into QueueTracks, while federated audio is synthesized as a remote queue track that streams its media URL directly. All embedded activity players share a module registry so starting one pauses the others and the global player, keeping a single audio source at a time |
frontend/src/components/statuses/ |
StatusComposer — shared status editor (plain text or Markdown, visibility, BCP-47 language defaulting to the browser locale, @ mention and # hashtag autocomplete (hashtags sorted by popularity; mentions also cover remote actors via the remote_users search flag and accept user@domain narrowing) plus track-only attach search via SearchBar/SearchSuggestions, file uploads through api/files.ts). Posts through api/statuses.ts (POST /statuses/) by default; the share dialog's Fediverse tab injects a custom submit that calls tracks.publishTrack instead, and ActivityEditModal reuses it for edits (initial* props seed the existing text/format/language/attachments; attachment chips map to songhive:fileId/songhive:trackId-marked docs). |
frontend/src/components/user/ |
Reusable user display components (UserLink) used across activity cards, resource owner metadata, file details, audit logs, and admin lists. UserLink renders local users as RouterLinks to /@{username}, remote users as external links to actor_url, and accepts either a full UserSummary owner or legacy username/displayName/avatarUrl/remoteUrl props |
frontend/src/components/admin/ |
Admin-specific shared components (e.g. StatCard for the dashboard) |
frontend/src/components/player/ |
Persistent player bar (PlayerBar, NowPlaying, QueuePanel, VolumeControl, OutputSelector) mounted in AppLayout so playback survives route changes, backed by stores/player.ts and the singleton player/engine.ts (dual HTMLAudioElement primary/preload). OutputSelector routes playback between "This device" and configured server-side outputs via stores/playback.ts; output management lives in the profile Outputs tab (views/OutputsView.vue, stores/outputs.ts). QueueTrack extends TrackResponse with stream_url — a direct media URL used instead of /api/v1/stream/{id} for audio without a local track row — and remote, which suppresses library links and listen-history reporting |
frontend/src/layouts/ |
App, auth, and admin layouts |
frontend/src/views/ |
Page-level components, including views/admin/ (Dashboard, Users, Settings, Reports, Invites, Audit, Tasks, Celery) behind the /admin guard (Home, Library, Album/Artist/Track/Playlist lists and details, History, Favorites, Files, File detail, Downloads, Radio station list/create/play, About, Login, Register, PasswordReset, VerifyEmail, /settings for the authenticated user, UserProfileView for /@{username}, UsersDirectoryView for /users, SearchView for the public /search page, plus 403/404 and placeholder views). SearchView renders grouped, independently sortable and paginated sections for every searchable entity via useSearchSections; the shared SearchBar supports an optional autocomplete mode backed by the aggregate /api/v1/search/ endpoint, with SearchSuggestions offering arrow-key navigation (Enter picks the highlighted item) via an exposed handleKeydown hook the controlling input forwards to. When the caller passes remote (the remote_available flag from the search response) and the query looks federated — an https:// URL or an @user@domain FQN — SearchSuggestions appends a "See on the Fediverse" entry that emits remote-lookup; SearchView routes it to /remote/lookup. UserProfileView renders user bios through the RichText component, which linkifies hashtags, mentions and URLs; library, playlist, album, and track detail views render the owner through the UserLink component |
frontend/src/api/ |
Typed HTTP client (openapi-typescript generated types.ts), per-resource modules including admin.ts for the admin panel, WebSocket event bus, stream URL helper |
frontend/src/i18n/ |
vue-i18n setup with lazy-loaded locales |
frontend/src/styles/tokens.css |
CSS custom properties for theming |
Browser authentication is cookie-based: the SPA never stores JWTs. The
api/client.ts fetch wrapper sends credentials: "same-origin" on every
request, echoes the readable csrf_token cookie as X-CSRF-Token on unsafe
methods, and retries once after a cookie refresh on 401. The auth store
(stores/auth.ts) persists only the non-sensitive user profile and bootstraps
by fetching /users/me; stream URLs (api/stream.ts) and the WebSocket
handshake (api/ws.ts) carry no token because the same-origin requests
authenticate through the HttpOnly cookies. This assumes same-origin
SPA/API hosting — which the production static serving and the Vite dev
proxy both provide — while non-browser clients keep using the token pair in
login/refresh JSON responses with Authorization: Bearer headers.
Entity detail pages expose an "Activities" action that navigates to
/{entity}/{id}/activities (EntityActivitiesView, shared across track,
album, artist, playlist, and library). The feed reads
GET /api/v1/{entity_type}/{entity_id}/activities with activity_type /
source_type filters and keyset cursor pagination; liking or boosting an activity
calls POST /api/v1/activities/{id}/like / /{id}/boost, replying calls
POST /api/v1/activities/{id}/reply, and editing one calls
PATCH /api/v1/activities/{id} (shown to the activity owner and admins).
ActivityCard renders a Mastodon-style action bar — reply, boost and like
icons with their counters — for authenticated users when the activity's
can_interact flag allows it (the icons act as toggles — liking/boosting
a reacted activity retracts it via DELETE /{id}/like / /{id}/boost;
clicking the like/boost
count opens ActivityActorsModal listing the known interactors from
GET /{id}/likes / /{id}/boosts, and clicking the reply count expands
the known replies from GET /{id}/replies — local replies rendered as
nested ActivityCards, remote ones as ActivityRemoteReply rows —
together with a StatusComposer wired to the reply endpoint). Card
content is never rendered as raw HTML — ActivityCard reduces it to safe
segments (text, line breaks, linkified mentions/hashtags/URLs) via
utils/activityContent.parseActivityContent; inline formatting elements
(<strong>, <em>, <code>, <del>, <u>, headings) are preserved as
marks on the segments and rendered through CSS classes rather than real
markup, and <ul>/<ol> items flatten to bullet/numbered lines with
indentation per nesting level. ActivityCard also renders
the activity's attachments (the AP attachment documents of the embedded
object): image media types inline, Audio/audio media types in an
<audio> player, everything else as a link. Reaction cards
(like/announce) are wrappers: ActivityObjectEmbed fetches the reacted
activity through a shared per-id cache (utils/activityFetch.ts) and
renders the full ActivityCard for Note objects or the compact item
card for Audio ones (falling back to an "unavailable" placeholder on
fetch failure), the card's timestamp/copy links point at the reaction's
object_url, and the Delete action retracts the caller's reaction
(stores/activities.retractReaction) rather than deleting the target.
UserProfileView (/@{username}) shows a Compose button to the profile
owner that opens components/statuses/StatusComposer.vue in a modal; the
share dialog's Fediverse tab reuses the same component for track
publication, and the default post format is configurable from
/settings (profile.status_content_type).
The home page (HomeView, components/home/) is a hybrid shelf + feed
layout that splits by audience and puts the two zones behind
Music | Activity AppTabs (shelves first) so the feed no longer trails
below the whole catalogue; both panels stay mounted (v-show) so tab
switches never refetch. Authenticated visitors get a greeting with
a StatusComposer modal, at most four "your music" shelves (Jump back in
from /history/ deduplicated by track, Favorites, Your uploads, New on
this instance), and the scoped activity feed. Anonymous visitors get the
instance hero (name, description, sign-in/register CTAs, optional
public_stats_enabled-gated counts from GET /api/v1/instance/stats),
public-catalogue shelves (recently added albums/tracks, public libraries,
people), genre chips, and the public instance feed. Every shelf fetches its
own data, renders a skeleton while loading, offers inline retry on error,
and removes itself entirely when empty. The feed (HomeFeed, backed by
stores/timeline.ts and GET /api/v1/timeline) offers a scope switch on
its own row — Mine | This instance | Federated for signed-in users,
This instance | Federated for anonymous visitors — and a
Posts | All activity mode switch on a second row. This instance only
surfaces locally sourced activities; Federated adds activities received
from remote instances and webmentions. Signed-in users default to Mine
when they have own activity, and the last explicit choice persists in
localStorage (a stored Mine is ignored once logged out).
The frontend is also a Progressive Web App. /manifest.webmanifest is served
from the backend so the PWA name follows the configured instance name and the
manifest theme_color/background_color react to the user's selected
light/dark theme. frontend/public/pwa/ contains generated icon variants
(including maskable-* icons for Android adaptive icons), and
frontend/public/sw.js provides a lightweight offline shell cache.
Build output is served as static files by the backend (or a CDN). The FastAPI
app sets router.default to an ASGI handler that serves files directly from
songhive/static/ and falls back to index.html for unhandled non-API paths,
so the Vue Router handles deep links such as /verify-email?token=....
When the fallback serves the SPA shell for an object page (/tracks/{id},
/albums/{id}, /artists/{id}, /playlists/{id}, /libraries/{id},
/genres/{name}, /tags/{name} and /@{username} plus their sub-pages),
api/semantic_meta.py injects semantic <head> tags into it: OpenGraph
og:title/og:description/og:url/og:type/og:site_name/og:image
metadata for social-media preview cards, and <link rel="tag"> elements for
the entity's hashtags. Ownership and related entities are expressed with the
OpenGraph music namespace and HTML authorship annotations: music:musician
links tracks and albums to their artist page, music:album links a track to
its album, music:creator links playlists/libraries to their owner, and
every uploadable entity also carries <link rel="author">, name="author"
and fediverse:creator tags pointing at the uploading user. Lookups go
through the regular ACL checks so private entities never leak metadata, and
any failure falls back to the unmodified shell.
Single-user mode is configured through the single_user_username admin
setting. When set, anonymous browser requests to / are answered with a
302 redirect to /@{username} by the always-mounted profile_pages
route, and the SPA router guard applies the same anonymous-only redirect
for client-side navigations; authenticated users keep the regular home
page and ActivityPub clients still get the SPA shell as before. Every
/ response in this mode also advertises rel="me" links to the
single-user profile — both the /@{username} and /users/{username}
forms, injected into the SPA <head> and the Link header (the header
is the only channel on the redirect, which has no body) — so link
verification (e.g. on Mastodon) succeeds even when only the instance
base URL is referenced. The setting is exposed as
single_user on /api/v1/instance and /api/v2/instance.
The Vite build also copies the swagger-ui-dist bundle into
songhive/static/swagger-ui/ and rewrites swagger-initializer.js to point
at the instance's own /openapi.json. FastAPI mounts those assets at
/swagger-ui/ (with a redirect from /swagger-ui), so interactive API docs
are available on every deployment without a separate container.
API Design
REST API under /api/v1/:
/api/v1/
├── auth/ # Login, register, token refresh, password reset, verify email, sessions
├── users/ # Public user profiles, directory, per-user activity feeds; authenticated profile updates
├── artists/ # Artist CRUD + search
├── albums/ # Album CRUD + search
├── tracks/ # Track CRUD + search
├── search/ # Lightweight aggregate multi-entity search (tracks, albums, artists,
│ # playlists, libraries, users, tags, genres) with normalized result
│ # items; ACL-filtered like the underlying list endpoints.
│ # A q starting with '#' is a hashtag lookup: the prefix is stripped
│ # and only the tags section is returned, sorted by item_count
│ # (a bare '#' lists the most used tags). With remote_users=1 the
│ # users section also matches remote actors cached by pubby
│ # (federation_followers + federation_actor_cache), returning
│ # 'user@domain' handles for mention completion
├── files/ # Generic file upload/list/download (StoredFile)
├── downloads/ # Bulk-download ZIP archives: create (explicit ids or
│ # album/artist/playlist/library container), list, status,
│ # file download, delete, clear completed; per-user active
│ # cap + rate limits, built asynchronously by Celery
├── libraries/ # Library management + add/remove tracks/albums/artists
├── playlists/ # Playlist CRUD + add/remove/reorder tracks/albums/artists + list tracks
├── favorites/ # Favorites/bookmarks
├── history/ # Listening history
├── radios/ # Dynamic radio generation
├── timeline/ # Cross-entity activity feed: scope=mine|instance (mine
│ # requires auth; anonymous defaults to instance),
│ # mode=posts|all with include_boosts/include_replies/
│ # source_type filters, keyset cursor pagination.
│ # No `following` scope — Songhive only receives follows
├── instance/ # Public instance metadata (Mastodon-compatible);
│ # `single_user` reports the single_user_username setting,
│ # /instance/stats exposes visibility-filtered counts when
│ # the public_stats_enabled admin setting is on (404 otherwise)
├── statuses/ # Standalone status posts (user-entity `create` activities with
│ # content type, language, file/track attachments, and mentions)
├── shares/ # Share grants (owner → specific user); /shares/mine lists the
│ # grants + tokens created by the current user (revoked
│ # tokens only with ?include_revoked=true)
├── share-urls/ # Share URL tokens (revocable short links)
├── share/{token} # Public short-URL resolver
├── reports/ # Content moderation reports (submit)
├── admin/ # Admin: settings, stats, user management, report review
└── stream/{id} # Audio streaming (Tornado handler, bypasses FastAPI)
Federation endpoints (mounted by pubby when federation is enabled):
/users/{username} # Per-user ActivityPub actor document (AP clients) or browser redirect to /@{username}
/users/{username}/objects/{id} # Dereferenceable ActivityPub objects (Audio, Note, Tombstone, stored payloads)
/users/{username}/quote_authorizations/{id} # FEP-044f QuoteAuthorization documents issued for the user actor
/@{username} # Mastodon-style profile: AP actor for AP clients, SPA shell for browsers with rel="me" links + OpenGraph tags
/tracks/{id} # Track page: Audio object for AP clients, SPA + rel=alternate hints + OpenGraph tags for browsers
/.well-known/webfinger # WebFinger discovery
/.well-known/nodeinfo # NodeInfo discovery document (pubby)
/nodeinfo/2.{0,1}[.json] # NodeInfo document (Songhive): pubby usage stats plus
# metadata.nodeName/nodeDescription, metadata.maintainer
# (configured contact person) and metadata.staffAccounts
# (actor URLs of active admins)
/ap/actor # Instance-level Application actor
/ap/inbox # Shared inbox (Songhive route; queues process_incoming)
/api/v1/*/ # Mastodon-compatible API (pubby adapter), except /api/v1/instance which is provided by Songhive and always available
Deployment
Docker Compose (docker-compose.yml) provides a reference deployment:
setup— one-shot root container that creates./volumessubdirectories and chowns them toPUID/PGID, so the other (rootless) services can write to themsonghive— application container (Tornado server)worker— Celery worker container (same image, different command)watcher— external-library filesystem watchdog (same image)stream-worker— server-side audio outputs (streamsprofile)postgres— PostgreSQL databaseredis— Redis (broker + cache + sessions)nginx— Reverse proxy (docker/nginx.conf). It proxies everything to the application, so the REST API, federation endpoints, frontend assets and the SPA shell are all served by FastAPI (api/app.pysetsrouter.defaultto a handler that servessonghive/static/files directly and falls back toindex.html). Serving the SPA through the backend lets it injectrel="me"links,rel="alternate"ActivityPub hints and OpenGraph/rel="tag"metadata into object pages (api/routes/profile_pages.pyandapi/semantic_meta.py), and perform the ActivityPub/browser content negotiation for routes that double as dereferenceable AP objects. Only two paths get special treatment:/ws/needs the WebSocket upgrade headers and/api/v1/stream/disables response buffering for real-time audio delivery.
Persistent data is stored under volumes/. config.toml sits next to
docker-compose.yml and is bind-mounted read-only at
/etc/songhive/config.toml; .env (see .env.example) holds PUID/PGID,
the host HTTP port and SONGHIVE_* overrides, which are injected into the
Songhive containers via env_file.
systemd
For non-Docker deployments, config/systemd/ contains unit files for running
Songhive as a systemd service. The install.sh script automates the setup:
- Creates a Python virtual environment
- Installs Songhive from the local checkout
- Copies
config.toml.exampleto the system (/etc/songhive) or user (~/.config/songhive) config directory - Creates data, cache, and log directories
- Installs the systemd units and sets the correct
ExecStartpaths
The master songhive.service unit starts four dependent units:
songhive-server.service— main web server (songhive)songhive-celery.service— Celery worker and schedulersonghive-watch-extlib.service— external-library watchdogsonghive-stream-worker.service— audio stream worker for server-side outputs (songhive stream-worker)