# Integrating SimplyBeautyKit with LiveKit SimplyBeautyKit plugs into LiveKit through the **VideoProcessor** of the local camera track. The room, signalling, encoding and transport stay as they are; the processor only changes the pixels of each captured frame before WebRTC encodes it. The local preview is fed after the processor, so it shows the effect too. ``` camera ─▶ LiveKit capturer ─▶ VideoProcessor ─▶ WebRTC source ─▶ encoder ─▶ room (SimplyBeauty) └──────▶ local preview ``` Two working samples implement this, one per platform. They build in this repository against the SDK sources and were run against a local LiveKit server: | | Android | iOS | |---|---|---| | Sample | [`examples/livekit-android`](../examples/livekit-android) | [`examples/livekit-ios`](../examples/livekit-ios) | | LiveKit SDK (pinned) | `io.livekit:livekit-android:2.29.0` | `client-sdk-swift` **exactly** `2.17.0` (SPM; pins LiveKitWebRTC 150.7871.2) | | Processor API | `io.livekit.android.room.track.video.NoDropVideoProcessor` (`livekit.org.webrtc.VideoProcessor`) | `LiveKit.VideoProcessor`: `process(frame:) -> VideoFrame?` | | Engine path | CPU I420 (`BeautyFrameHook`, `SimplyBeautyEngine.process`) | CPU `CVPixelBuffer` (`-processPixelBuffer:output:rotation:mirror:frameType:`) | | Toolchain | Gradle 8.11.1 wrapper, AGP 8.7.2, Kotlin 2.0.21, JDK 17+ (built with 21), compileSdk 35, minSdk 24 | Xcode 16.3+ (built with 26.3), XcodeGen, iOS 15+, Swift 5 mode | | SDK dependency | `com.simplyapphub:simplybeauty` + `:simplybeauty-models`, substituted with `../../android` `:lib` / `:models` (composite build) | local package `../..` (`SimplyBeauty`, bundled models) | | Verified on | API 34 arm64 emulator: 4 instrumented + 10 JVM tests; published to a local server | iOS 26.3 Simulator (iPhone 16e): 10 XCTests; published to a local server | The processor source files are the reference implementation: copy `BeautyVideoProcessor` (+ `EngineBeautifier` on Android) into your app. The snippets below are taken from them and compile there. ## Android ### Dependencies ```kotlin // settings.gradle.kts: livekit-android needs JitPack for one dependency // (com.github.davidliu:audioswitch) dependencyResolutionManagement { repositories { google() mavenCentral() maven("https://jitpack.io") { content { includeGroup("com.github.davidliu") } } } } // app/build.gradle.kts dependencies { implementation("com.simplyapphub:simplybeauty:0.2.0") implementation("com.simplyapphub:simplybeauty-models:0.2.0") // face tracker + masks implementation("io.livekit:livekit-android:2.29.0") } ``` The core library ships no models (size criterion C6); `simplybeauty-models` carries them, and `SimplyBeautyModels.install(context)` copies them once to app storage (call it off the main thread) and returns the `modelDir` / `resourcePath` for the engine config. ### Wiring ```kotlin val paths = withContext(Dispatchers.IO) { SimplyBeautyModels.install(applicationContext) } val engine = SimplyBeautyEngine() val beautifier = EngineBeautifier(engine, { w, h -> BeautyLook.config(paths, w, h) }, BeautyLook::apply) val processor = BeautyVideoProcessor(beautifier) room.connect(url, token) val track = room.localParticipant.createVideoTrack( name = "camera", options = LocalVideoTrackOptions(position = CameraPosition.FRONT), videoProcessor = processor, ) track.startCapture() track.addRenderer(localView) // processed preview room.localParticipant.publishVideoTrack(track) processor.enabled = false // beauty off: frames pass through untouched // teardown: stop the camera first, then close the processor (waits for the // frame in flight, then releases the engine) track.stopCapture() room.localParticipant.unpublishTrack(track) track.dispose() processor.close() ``` The track has to be created with the processor: `setCameraEnabled(true)` and Compose `RoomScope` create their own camera track without one. ### What BeautyVideoProcessor does - **Capture thread** (WebRTC's `SurfaceTextureHelper` thread): converts the frame to I420 with `frame.buffer.toI420()`. For camera texture frames that is WebRTC's GPU read-back on that same thread, so the camera texture is free again when `onFrameCaptured` returns; the processor never keeps the camera frame. - **One worker thread** runs the engine. Between the two sits a one-frame slot: when the engine is slower than the camera, the waiting frame is replaced by the newer one (`droppedFrames`), so the capture thread never waits for the effect and latency stays one frame. - **Output**: a `JavaI420Buffer` (native memory, freed on release) with the input's size, rotation and timestamp. - **Rotation**: camera buffers are in sensor orientation and `VideoFrame.rotation` is metadata. The engine gets it (`BeautyFrameHook.processI420(..., mirror, rotation)`), so face tracking sees an upright face, and writes the result in the buffer's orientation; the frame keeps its rotation. The published stream is not mirrored (`MirrorMode.NONE`); LiveKit's renderer mirrors the front-camera preview. - **Size changes**: WebRTC adapts the capture size at run time (CPU and bandwidth adaptation, camera switches). `BeautyFrameHook` resizes the existing engine in place, retaining loaded models, settings, statistics and the quality governor. Injected faces and masks are cleared on an actual size change; inject fresh data for the new dimensions. - **Errors**: a non-zero status or an exception sends the frame unprocessed (`failedFrames`, `lastStatus`). - **Timestamps** sent to the sink strictly increase: switching `enabled` off while the worker is busy would otherwise send an older processed frame after a newer camera frame (WebRTC drops those anyway); the processor drops it itself (`staleFrames`). GPU alternative (not used by the sample, and not run under LiveKit yet): camera frames arrive as OES `TextureBuffer`s, and the engine's GLES path (`Format.GPU`, `glInit` / `processTexture`) can render them without the read-back. See [GPU texture path](#gpu-texture-path-android) below. ### Local development server Debug builds of the sample allow `ws://` to `10.0.2.2` and `localhost` only (`app/src/debug/res/xml/network_security_config.xml`); Android blocks cleartext otherwise (`CLEARTEXT communication ... not permitted`). From the emulator, signalling goes to the host at `10.0.2.2`, but media must go to the host's LAN address (ICE to `10.0.2.2` failed in our runs): ```sh livekit-server --dev --bind 127.0.0.1 --bind --node-ip lk token create --dev --join --room demo --identity android --valid-for 1h --token-only cd examples/livekit-android ./gradlew :app:installDebug -Plivekit.url=ws://10.0.2.2:7880 -Plivekit.token= ``` The URL and token can also come from `livekit.url` / `livekit.token` in `local.properties` (git-ignored) or `LIVEKIT_URL` / `LIVEKIT_TOKEN`; they end up in `BuildConfig` as the defaults of the app's fields. Never commit a token. ## iOS ### Dependencies ```yaml # XcodeGen (examples/livekit-ios/project.yml) packages: LiveKit: url: https://github.com/livekit/client-sdk-swift exactVersion: 2.17.0 SimplyBeauty: path: ../.. # or the SimplyBeauty package URL and a version ``` LiveKit 2.17.0 declares `swift-tools-version:6.1` (Xcode 16.3 or later). The models come from the SimplyBeauty package's resource bundle (`SBConfig.useBundledResources = true`). ### Wiring ```swift let processor = BeautyVideoProcessor() // keep a strong reference: LiveKit holds it weakly try await room.connect(url: url, token: token) let track = await LocalVideoTrack.createCameraTrack( options: CameraCaptureOptions(position: .front), processor: processor) try await room.localParticipant.publish(videoTrack: track) // SwiftUIVideoView(track) shows the processed frames. processor.isEnabled = false // beauty off: frames pass through untouched ``` The processor can also be set on an existing track (`localVideoTrack.processor = processor`), e.g. on a buffer track: the sample publishes a generated test pattern that way when there is no camera (the Simulator). ### What BeautyVideoProcessor does - LiveKit calls `process(frame:)` on the capturer's serial processing queue and skips new camera frames while one is running, so the processor keeps no queue of its own. - **Output**: a buffer from a `CVPixelBufferPool` in the input's pixel format (the camera delivers 420f; 420v and BGRA work too), IOSurface-backed, with the frame's dimensions, rotation and timestamp. Other formats (e.g. 32ARGB from a buffer track) pass through. - **Rotation**: `frame.rotation.rawValue` (clockwise degrees) goes to the engine; the output stays in the buffer's orientation. - **Adapted frames**: when the camera format is larger than the requested capture size, LiveKit crops and scales lazily, so `frame.dimensions` is smaller than the pixel buffer. The processor renders the adapted image (centre crop, then scale, as WebRTC does) into a BGRA buffer first. - **Size changes**: `resize(width:height:)` updates the existing engine, retaining its models, current look and quality governor. A size whose engine cannot be created or resized is not retried every frame; a failed resize keeps the previous engine usable at its previous dimensions. - **Errors**: the frame is returned unprocessed (`stats.failed`, `stats.lastStatus`). - **Debug builds are slow**: SwiftPM compiles the C++ core with the app's configuration, so at Debug (-O0) the CPU pipeline is 5-50x slower. The sample's scheme runs Release. ### Local development server ```sh livekit-server --dev cd examples/livekit-ios && xcodegen generate xcodebuild -project SimplyBeautyLiveKit.xcodeproj -scheme SimplyBeautyLiveKit -configuration Release \ -destination 'generic/platform=iOS Simulator' -derivedDataPath build/dd build TOKEN=$(lk token create --dev --join --room demo --identity ios --valid-for 1h --token-only) xcrun simctl install booted build/dd/Build/Products/Release-iphonesimulator/SimplyBeautyLiveKit.app xcrun simctl launch booted com.simplybeauty.sample.livekit \ -LiveKitURL ws://localhost:7880 -LiveKitToken "$TOKEN" -LiveKitAutoConnect YES ``` Launch arguments win over the `LIVEKIT_URL` / `LIVEKIT_TOKEN` build settings (`Config.xcconfig`, overridable in the git-ignored `Config.local.xcconfig`). `ws://` is allowed for local-network servers only (`NSAllowsLocalNetworking`). ## Build and test commands (verified) Run from the repository root. Android device commands need a device or emulator; with several attached, choose one with `ANDROID_SERIAL`. ```sh # Android sample cd examples/livekit-android echo "sdk.dir=$HOME/Library/Android/sdk" > local.properties ./gradlew --no-daemon :app:assembleDebug # BUILD SUCCESSFUL ./gradlew --no-daemon :app:testDebugUnitTest # 10 tests, 0 failures adb push portrait.jpg /data/local/tmp/sb_portrait.jpg # optional: face test ANDROID_SERIAL=emulator-5672 ./gradlew --no-daemon :app:connectedDebugAndroidTest # 4 tests, 0 failures, 0 skipped # iOS sample cd examples/livekit-ios xcodegen generate xcodebuild -project SimplyBeautyLiveKit.xcodeproj -scheme SimplyBeautyLiveKit \ -destination 'generic/platform=iOS Simulator' -derivedDataPath build/dd build # BUILD SUCCEEDED TEST_RUNNER_SB_TEST_IMAGES= xcodebuild -project SimplyBeautyLiveKit.xcodeproj \ -scheme SimplyBeautyLiveKit -destination 'platform=iOS Simulator,name=iPhone 16e' \ -derivedDataPath build/dd test # 10 tests, 0 failures ``` What the tests cover: - Android JVM (`BeautyVideoProcessorTest`, fake effect, WebRTC frames over direct buffers): output keeps size, rotation and timestamp; the conversion runs on the capture thread and the camera frame is not kept; the waiting frame is replaced by the newest; errors and exceptions send the frame unprocessed and the worker survives; off sends the captured frame itself; sent timestamps never go back; a new capture session resets that; output follows size changes; `close()` releases every buffer and the effect. - Android device (`BeautyVideoProcessorDeviceTest`, real engine + installed models, `JavaI420Buffer.allocate`): 640x480 frames at rotation 90 are processed and changed; 640x480 -> 320x240 -> 1280x720 retains the engine frame count and a live edit to the look; off = pass-through with no engine; on a face photo delivered as a sensor buffer the tracker reports level eyes only with rotation 90 (with rotation 0 it reports the face on its side). - iOS (`BeautyVideoProcessorTests`, real engine + bundled models): BGRA, 420f and 420v stay in their format with dimensions, rotation and timestamp; off returns the same frame; size changes retain engine identity and live edits; adapted 1280x960 buffer -> 640x360 frame; unsupported format passes through; a failing engine is not retried per frame; the processor inside LiveKit's own buffer-capturer pipeline renders its output; the rotated-face test as on Android; launch arguments win over Info.plist settings. - Each test fails when the behaviour it covers is removed: 9 single-point changes of the Android processor, 3 of the Android engine adapter and 8 of the iOS processor were each caught. End-to-end runs against `livekit-server` 1.13.6 (`lk` 2.18.6): - Android, API 34 arm64 emulator (emulated back camera, no front camera, so the sample falls back to the back one): the participant published a 1280x720 VP8 simulcast camera track (320x180 / 640x360 / 1280x720 layers); the processor sent the emulator camera's frames (about 9 fps; a few were dropped while the engine started or ran slowly) with 0 failures; the quality governor stepped HIGH -> MEDIUM -> LOW on the emulator's slow CPU; WebRTC's adaptation changed the frames to 960x540 and the engine was re-created; the beauty switch stopped and resumed processing; disconnect and reconnect worked. - iOS Simulator: the test pattern (720x1280 BGRA, 15 fps) through the processor, published as a 720x1280 VP8 simulcast camera track; 463 frames processed, 0 failed, about 21 ms per frame at MEDIUM in Release. Emulator and Simulator timings are not performance evidence. On a Galaxy Z Fold7 the engine's production path at 720p (built-in tracker + segmenter + full beauty stack) takes 8.03 ms p50, 11.46 ms p95 (engine benchmark outside LiveKit, see [PERFORMANCE.md](PERFORMANCE.md)); a LiveKit device run is still to be done (C1). ## Frame-format checklist - Keep `rotation` and the timestamp of every frame. Tell the engine the rotation; do not rotate the pixels. - Android camera frames are OES `TextureBuffer`s: `toI420()` on the capture thread (CPU path) or the GPU path below. Do not keep camera frames across threads: the camera has a small number of textures. - Never block the capture thread on the effect: one worker, newest-frame slot (Android), or LiveKit's own frame skipping (iOS). - Send monotonically increasing timestamps. - Resize the existing engine when the buffer size changes (Android `resize()` / `BeautyFrameHook`, iOS `resize(width:height:)`). Re-inject faces and masks after an actual size change if you inject them. - Local preview: add the renderer to the processed track; mirror the front-camera preview in the renderer, not in the published pixels. - Multi-face / pose / expressions: call `injectFaces` from your platform landmarker before processing, then `getFacePose()` / `getFaceExpressions()` after (or use the built-in tracker as the samples do). ## GPU texture path (Android) Not used by the sample; the steps for rendering WebRTC texture frames on the GPU instead of `toI420()` + `processI420`: 1. Own **one GL thread** ("BeautyGL") with an ES 3 pbuffer context shared with the app's **single, process-lifetime root `EglBase`** (the encoder and renderers share that group, so they can read our output textures; `LiveKitOverrides(eglBase = ...)` passes it to LiveKit). 2. On that thread create one `Format.GPU` engine for the life of the process and call `glInit()` once (shader compile); `glResize(w, h)` whenever the adapted buffer size changes (cheap, no recompile). Do both as **posted** tasks, not inside the per-frame hop; pass frames through until ready. 3. Per frame (GL thread): draw the OES buffer into an RGBA8 `src` texture with `FLIP = T(.5,.5)·S(1,-1)·T(-.5,-.5)` so storage is top-down, inject the face (buffer-space landmarks, weight in (0, 1]), then `processTexture(src, dst, w, h, frame.timestampNs)` into a pooled RGBA8 `dst`, `glFinish()`, and wrap `dst` in a `TextureBufferImpl` (type RGB, transform `FLIP`). Forward `VideoFrame(out, frame.rotation, frame.timestampNs)`. 4. GL state after `processTexture`: FBO 0, VAO 0, `ARRAY_BUFFER` 0, `PIXEL_PACK`/`PIXEL_UNPACK` buffers 0, program 0, `UNPACK_ALIGNMENT` 4 — what WebRTC's client-side-array drawers and `YuvConverter` need. Output alpha is 1. 5. Never make GL calls in a `TextureBuffer` release callback (it may run on the encoder or renderer thread): post back to the GL thread. 6. Camera session end: `glTrim()` (then `processTexture` returns 2 until the next posted `glResize`/`glInit`; pass frames through meanwhile). Teardown: `glRelease()` on the GL thread while the same context is current, then `release()` on that thread (never while a frame is running). After a context loss, `glRelease()` on the old context if it is still usable, otherwise just `glInit()` on the new one (the old objects are dropped). A GPU engine never creates the face tracker or segmenter: inject faces. Coordinate conventions: - Buffer space is `buffer.width x buffer.height`, sensor orientation, not mirrored; normalised `(u, v)` has its origin at the top left. - `frame.rotation` (clockwise turn that makes the buffer upright) is metadata only; keep it on the forwarded frame. - Upright detector coordinates `(X, Y)` map to buffer `(u, v)` by rotation: 0: `(X, Y)`; 90: `(Y, 1-X)`; 180: `(1-X, 1-Y)`; 270: `(1-Y, X)`. - No x-flip for the front camera; the face masks use symmetric landmark sets. ## Performance budget | Stage | Target @720p | |---|---| | Face mesh | ≤ 8 ms | | Effects (all on) | ≤ 4 ms | | Frame conversion (Android texture → I420 read-back) | ≤ 2 ms | | Total overhead | ≤ 16 ms (criterion C1, flagship), leaving headroom at 30 fps | The samples use `Quality.AUTO` with `setFrameBudgetMs(30)`: when frames take longer, the governor steps HIGH → MEDIUM (no makeup, teeth or eye brightening; tracker every other frame) → LOW (cheap smoothing and whitening only), and back up when there is headroom (`QUALITY_CHANGED` events, `Stats.activeQuality`).