The RAI Convention Centre in Amsterdam, Netherlands

IBC 2026 Technical Conference Highlights

Hero image courtesy Deposit Photos

David Kirk reports from the 11-14 September 2026 International Broadcasting Convention

Held in mid September at the RAI Amsterdam convention centre, the International Broadcasting Convention is Europe’s longest established exhibition of media technology hardware, software and services. Attendance was slightly down this year and exhibitor numbers indicate considerable churn but the event as a whole retained its focus and sparkle. The following report is a selection of IBC technical conference presentations highlighting the issues currently of central interest to broadcasters. Don’t expect an easy read; the technology is advancing fast.

Unified broadcast control

Georg Fürst (Austrian Broadcasting Corporation, ORF) promoted the advantages of unified broadcast control – From heterogeneous record and playout systems to a unified control room solution at ORF:

“In modern broadcast environments, efficient and adaptable control room systems are essential. Traditional multi-system architectures can raise operational complexities and service demands, while the reduction in staff necessitates solutions where one operator can efficiently oversee multiple playout and recording tasks concurrently. A novel solution implemented by ORF unifies file management, recording, and contribution playout within a single application, aimed at minimising interface complexity and optimising operations.

“The selected approach involves renewing and consolidating video server infrastructures and establishing a uniform device type, which simplifies maintenance, reduces support requirements, and enhances cross-departmental expertise. Consequently, the new architecture sets a standard for ORF’s technical infrastructure, strengthening both the robustness and scalability of TV production, facilitating sustainable high-quality content production, and improving organisational resilience amid shifting demands. Successful rollout and interest from other broadcasters suggest promising prospects for broader adoption as a benchmark model in future control room environments.

“From the outset, the project team pursued the guiding principle of minimising the number of controls visible during live operation while ensuring that all necessary functions remained accessible. Accordingly, an early design decision was made to provide extended configuration options through collapsible tabs and right-click context menus. This approach ensures that operators are presented with only the essential controls during a production, while retaining full access to the complete functional scope within a single application during pre- and post-production phases.

“The Player offers a robust interface with tabs like Player Window, Playlist (Cliplist), Recorder and Marker, allowing users comprehensive control over video and audio playback. Essential operations include starting playback, managing playlist associated with the channel (Cliplist), recording, with functionalities like synchronous playback and clip customisation through subclip creation.

“The Window Channel consists of a lot of features which can be opened by the operator as needed. It ensures that operators are presented with only the essential controls during a production, while retaining full access to the complete functional scope if needed. See Figure 1. The WebRTC Player provides a preview of the injected clip and is showing the channel output of the server. Other functions are to open the Media Base View (with the file directory) to access and organise files or to show the remaining disk space on the server.

“The new architecture establishes a benchmark for the technical infrastructure of ORF control rooms and enhances both the robustness and scalability of daily television production.”

AoIP-based audio production

T. Onishi and M. Okano (Japan Broadcasting Corporation, NHK) addressed the subject of Feasibility and deployment strategies for cloud-based AoIP audio consoles.

“Demand for flexible and resource-efficient broadcast production continues to increase, particularly in facilities with constrained technical resources. IP-based and distributed production workflows have been widely discussed as alternatives to traditional facility-centric architectures, and cloud-based production models have emerged as potential options for selected workflows. However, migrating broadcast-grade audio systems to cloud environments presents significant technical challenges.

“Public-cloud environments exhibit characteristics that complicate the direct application of conventional AoIP architectures. Network traffic paths are dynamically managed to support load balancing and fault tolerance, they cannot be treated as static or topology-specific. Packet delay variation is influenced by multi-user resource contention rather than nominal bandwidth alone. In addition, native IP multicast routing, assumed by professional AoIP systems, is not inherently supported in public-cloud networks and typically requires additional service components or architectural adaptations, which tend to be expensive to introduce.

“Two principal approaches can be considered for synchronising time across geographically separated sites: 1. Deploying a PTP Grandmaster (GM) at each site and operating independent PTP domains, or 2. Extending a single PTP domain by transporting PTP from one site to the other using a single GM.

“While the second approach can be effective in dedicated network environments, transporting PTP across networks with unstable jitter characteristics (such as 5G systems or best-effort public networks) remains challenging. The lack of multicast support in public-cloud environments further limits the practicality of this approach, even with dedicated connectivity. Applying AoIP standards directly under such conditions can lead to unpredictable timing behavior, inconsistent buffering, and potential loss of phase coherence.

“Broadcast PTP is based on International Atomic Time (TAI), whereas system time provided within public-cloud environments is based on Coordinated Universal Time (UTC). Because cloud environments rely on UTC-based hardware clocks rather than a TAI-based PTP time, explicit buffer management becomes critical to compensate for timing drift between domains. However, sample-accurate alignment is not necessary for mixing operations that do not reference an external clock.

Figure 2 illustrates the overall cloud-based AoIP system architecture with separated transport and timing domains. A proof-of-concept implementation was conducted to validate the practical feasibility of combining a unicast-based AoIP transport and hybrid timing operation. This evaluation used a single 32-channel uncompressed PCM audio stream generated by on-premises AoIP input/output equipment. Audio was distributed within the on-premises network using multicast transport based on AES67/SMPTE ST 2110-30, converted to unicast at the cloud boundary, and transported to a cloud-based audio console hosted on a virtualised compute instance. The internal processing latency of the evaluated audio console was 128 samples.

“Experimental results demonstrate that cloud-based AoIP systems are most effectively deployed as incremental functional extensions rather than full replacements of existing broadcast audio infrastructures.”

Cost-efficient virtual production

Osmia VP, a low-cost, field-deployable virtual production system for in-camera VFX was outlined by H. Raymond-Hayling and colleagues (British Broadcasting Corporation Research & Development):

“We seek to reduce the barrier to entry for virtual production, which relies on specialist equipment and proprietary software and is therefore too costly for mid- to low-budget productions. These productions make up a significant portion of the 20,000 hours of content the BBC produces annually, in line with its public service remit. We will demonstrate the system and its integration of scene capture with real-time tracking and projection; all on commodity hardware. Developed in consultation with BBC production teams, the system enables the decoupling of location capture and performance capture, allowing realistic virtual backgrounds captured via Gaussian splatting to be combined with live-action footage filmed in studio.

“The technical architecture is built on open-source foundations. Osmia (Figure 3) employs Gaussian splatting for virtual environments, a modified SLAM approach (AsLAM –Asynchronous Location and Mapping) with printed markers for real-time camera tracking, and a modified Gaussian splat renderer for perspective-correct projection. This workflow could significantly reduce the cost and environmental impact of location filming and television production more generally. A scene is captured by a single camera operator with a camera or mobile phone. A Gaussian splat is trained on stills or a video using cloud compute and the Gsplat implementation. We used an NVIDIA A100 which can train a scene of 400 images in about 2 hours.

“Accurate camera tracking is essential for virtual production, ensuring virtual backgrounds align correctly with live action footage. Conventional systems are prohibitively expensive, requiring dedicated spaces and high end optical tracking systems. Our system requires no specialist equipment and can be deployed ad hoc in any studio space, including non specialist environments such as offices.

“The most significant limitation is scale: a single consumer screen constrains shot size and number of performers. A potential solution which would not incur greater costs could be compositing footage from multiple smaller screens. Wider adoption will also require sustained knowledge sharing on how production techniques and planning can evolve alongside virtual production.

“Since this production workflow is intended for use in situations where rapid deployment is desirable, it would be beneficial to automate as much of the process as possible. Therefore, it would be very useful to develop a system to colour grade the background to match the foreground. The foreground lighting may be highly constrained due to the circumstances of filming, but the colour of the background may be changed more easily.

“It would also be useful to introduce moving elements into the background for greater realism and creative possibility. Capturing a full 4D Gaussian Splat may be a challenge to do quickly. However, it is possible that a small number of 2D video elements could be placed inside a larger 3D scene to add natural movement.

“It would also be useful to be able to add some foreground Gaussian splats as well. Depending on the scene, some amount of foreground, such as leaves in a jungle setting, could be dynamically composited on top of the video output of the main camera. We believe Osmia VP demonstrates that virtual production need not require specialist infrastructure to achieve a high quality output, and that this approach has potential to reduce both the cost and environmental impact of television production.”

Scene-adaptive cameras

K. Tomioka and colleagues (NHK, Japan) with S. Kawahito (Shizuoka University) spoke on the theme of Scene-adaptive camera: dynamic imaging parameter control for localised shooting areas.

“Achieving high spatial resolution, high frame rate, and high dynamic range simultaneously remains challenging for imaging systems because of the inherent trade-offs among readout speed, power consumption, and data rate limitations. To overcomes these limitations, we propose a scene-adaptive camera that, instead of applying uniform settings across entire frames, optimises imaging parameters for individual regions within each image. The proposed approach enables localised control of spatial resolution, frame rate, and exposure time at the fine granularity of 4 × 4-pixel blocks. To demonstrate the feasibility of this concept, we developed a camera system based on a novel 4K CMOS image sensor capable of independently assigning one of four imaging modes to each block. The camera also incorporates real-time feedback control that dynamically analyses local brightness and motion within the scene and updates the parameters in synchrony with sensor operation.

“Figure 4 shows the developed CMOS image sensor. The sensor incorporates an effective pixel array of 3,904 (H) × 2,224 (V) pixels, corresponding to approximately 8.7 megapixels. The array is divided into 4 × 4-pixel control blocks, yielding a total of 976 (H) × 556 (V) blocks. Four imaging modes (Normal, Bright, Low-Light, and Fast) can be assigned independently to each block according to the characteristics of the subject. The developed 8.7 megapixel CMOS image sensor achieves highly fine-grained control at the level of 4 × 4 pixel blocks through a novel architecture in which the pass-gate circuits are integrated within the column circuitry.

“This architecture enables precise local adaptation to subject contours and complex brightness distributions, allowing spatial resolution, frame rate, and exposure time to be controlled independently within a single frame. The sensor was integrated with an RGB three-chip optical system and a real-time feedback control system to create a practical camera platform. The resulting system simultaneously achieved a wide dynamic range and a frame rate of 240 fps while reducing motion blur and maintaining 4K video quality. Extending the underlying elemental technologies to a practical 4K system suitable for broadcast production represents an important milestone in this study.

“These results demonstrate the effectiveness of locally allocating imaging resources according to scene characteristics. The proposed technology provides a foundation for next- generation imaging systems that maximise image quality while limiting data rate requirements, including applications in immersive media. Future work will focus on further enhancing the scene analysis algorithms. In addition, by exploiting the unique capabilities of the proposed approach, we aim to deploy this technology in a wide range of applications, from broadcasting to broader social infrastructure systems.

“Experiments show that the proposed camera substantially improves image quality, simultaneously suppressing overexposure and underexposure while reducing motion blur within a single frame.”

Converging IP-based systems and AI

MXL as an interoperability boundary between artificial intelligence and broadcast was the theme of a presentation by Guillaume Arthuis and colleagues (BBright, France) with S. Moubayed and colleagues (Hexalia, France).

“This paper presents an architectural approach based on the Media eXchange Layer, an open standard promoted by the European Broadcasting Union, which establishes a clear interoperability boundary between broadcast infrastructures and AI processing systems. The core hypothesis is that AI integration challenges are best addressed by introducing a standardised abstraction layer separating media transport from media intelligence.

“Figure 5 shows the structure of an AI processing module designed to consume and produce synchronised media MXL interfaces while separating preprocessing, inference, rendering and supervision. This separation of concerns allows AI engines to operate independently of underlying infrastructures, reducing workflow-specific integration effort. The AI module is not limited to inference. It includes four functional stages: preprocessing, inference, result interpretation and rendering or metadata generation.

“Preprocessing adapts the broadcast signal to the representation required by the model. Inference produces model outputs, usually detection metadata such as zones, labels and confidence values. Result interpretation converts model outputs into workflow actions. Rendering or metadata generation applies those actions to the outgoing media or exposes them to downstream systems. This distinction is important because many AI models produce information that is not directly usable in a broadcast chain.

“Early implementation feedback demonstrates reduced integration complexity, improved scalability, and stronger interoperability compared with traditional point-to-point AI integrations. AI adoption in broadcast is constrained less by the absence of algorithms than by the lack of a stable boundary between real-time media infrastructure and AI processing.

“Broadcast requires deterministic timing, signal integrity and operational continuity; AI requires flexible runtimes, model evolution and accelerated compute. Direct coupling leads to bespoke and fragile integrations. This paper proposed an MXL-based architecture in which the Media eXchange Layer acts as that boundary. The broadcast domain remains responsible for ST 2110 media handling, synchronisation and continuity, while AI modules operate as containerised services connected through explicit data and control planes.

“We conclude that industrialising AI in broadcast depends less on algorithmic innovation than on architectural standardisation. By defining a stable boundary between AI and broadcast domains, MXL enables broadcasters to combine broadcast-grade reliability with the innovation pace of AI technologies.”

LED-based virtual production

C. Borowski, Südwestrundfunk, Germany, explained the merits of A 100 Hz frame-interleaved approach to live multicamera switching in LED-based virtual production.

“SWR, the public broadcaster for southwestern Germany, is preparing a renewal of its Baden-Baden studio infrastructure. As part of this preparation, SWR ran a virtual production proof of concept to test, under real production conditions rather than in a closed lab, whether and how LED-based virtual production can support the formats that future studios at the site are expected to host. The proof of concept ran as a three-month evaluation phase between September and November 2025, with most of the technical system rented for the duration. SWR has main sites in Baden-Baden, Stuttgart and Mainz with different production profiles: Stuttgart and Mainz host most of the news and high-rotation studio output, with main-stage facilities that turn over multiple productions per day and set changes of around 15 minutes. Baden-Baden has its emphasis on entertainment, cultural and factual programming; the studio infrastructure there is reaching the end of its operational life and a substantial renewal is anticipated.

“Three questions framed the proof of concept:

“Q1: Production fit: Can LED-based virtual production be made to work under live multicamera broadcast conditions, given that published experience originates from single-camera, non-live film work?

“Q2: Site fit: Where within SWR would an investment be productive, given the difference in studio utilisation across sites?

“Q3: Format fit: Under what conditions is Virtual production genuinely needed by a format, rather than added as a visual upgrade?

“The central case was an eight-hour live pen-and-paper format produced for the ARD Twitch channel, with three tracked cameras and live cutting on a 10 by 4 m LED wall driven at 100 Hz. The technical contribution is a frame-interleaved approach for live multicamera switching (Figure 6). Odd frames on the wall carry the perspective-correct render of the active camera, even frames carry a uniform chroma-key fill. Each camera has its own render node running continuously, so the live cut is a phase change on the wall, not a switch in the render path. Inactive cameras’ previews are assembled by chroma keying against parallel reduced-quality renders.

“The proof of concept allows the three framing questions to be answered with a level of confidence that was not available before.

“On Q1 (production fit), LED-based virtual production can be operated under live multicamera broadcast conditions provided a method addresses the multicamera constraint directly; the frame-interleaved approach did so across eight hours of live production.

“On Q2 (site fit), the relevant constraint is not daily operation but format development, which requires contiguous studio time at the actual render and LED hardware. For SWR, this implies that the question is less which site can host operation than where the development capacity should sit; the Baden-Baden site, with its longer-form output and the upcoming studio renewal, remains the candidate where development and operation can sensibly co-exist.

“On Q3 (format fit), two findings are organisational rather than technical and apply to any public-service broadcaster considering a similar investment. The first finding is the cost shift toward a dedicated graphics capability: a virtual production setup is only as productive as the team that produces and maintains its content, and the savings on traditional scenery only materialise if such a team exists, involved from the concept stage rather than only from execution. The second finding is the role of continuous internal communication: editorial staff cannot generate format ideas for a technology they have never seen at work, and a future investment should include a deliberate communication and editorial onboarding component as part of the project, not as an afterthought.”

AI-assisted editing

Sovereign AI-assisted editing: MCP versus agent-to-agent orchestration for confidential audiovisual content was summarised by A. Rouxel and P. Fouché (European Broadcasting Union, Switzerland), A. Messina (Radiotelevisione Italiana), and R. Salomone (SRG SSR, Swiss Broadcasting Corporation, Switzerland)

“Broadcasters increasingly want AI-assisted editing while keeping unpublished video and audio inside infrastructure they control. This paper presents a metadata-centric, local-first architecture in which raw media is processed locally and editorial agents receive only time-aligned metadata, such as transcripts, shot boundaries, face clusters, scene descriptions, and acoustic events. The architecture has three layers: intent and planning, editorial reasoning agents, and Model Context Protocol (MCP)-based media-analysis services. An optional local de-identification gateway removes or replaces configured identifiers before selected metadata is sent to an external reasoning model. We compare two ways of coordinating the same agents and services: a centrally planned MCP-based workflow and collaborative Agent-to-Agent (A2A) coordination. The comparison is qualitative and focuses on control, reproducibility, flexibility, auditability, confidentiality, and operational cost. We propose a hybrid approach: central orchestration for fixed, repeatable tasks and A2A collaboration for decisions that depend on an underspecified editorial brief.

“Newsrooms face a tension between the rapid adoption of generative and multimodal artificial intelligence and the confidentiality, governance, and rights obligations attached to broadcast content. In this paper, sovereignty means that a broadcaster retains control over where raw content is stored, processed, and accessed. Unpublished footage may include embargoed information, protected sources, identifiable people, or rights-restricted archive and sports material. Sending such video or audio to a third-party AI service can therefore expose sensitive editorial material and create data-protection, ownership, and compliance risks.

“AI assistance can nevertheless add value to editorial and archive workflows. Prior work on facial recognition for television content showed how time-aligned, user-oriented metadata can support media professionals. MCP and A2A provide complementary standards for connecting AI systems to tools and for enabling agents to communicate. Together, these developments make it possible to consider partially automated workflows that segment, rank, and assemble content under a journalist’s high-level instructions.

“This paper proposes a local-first architecture that keeps raw audiovisual assets inside the broadcaster’s controlled environment while allowing AI systems to reason over derived metadata. The architecture is organised from the user downwards: Layer 1 interprets the editorial request and sets constraints; Layer 2 contains specialised agents that make editorial decisions from metadata; and Layer 3 contains local media-analysis services that create time-aligned descriptions of the source material.

“Future work will compare three deployment profiles: fully local, with metadata and language models inside the broadcaster’s perimeter; hybrid, with local metadata and an external language model receiving a de-identified view; and fully external, with a third-party multimodal model accessing raw content. The third profile provides a baseline for measuring the performance and sovereignty trade-off of externalisation. Another direction is integration with the BBC Time-addressable Media Store (TAMS) (7), allowing MCP data to reference an established time-based media store while keeping the typed-metadata layer unchanged.”

IBC 2027

Real trade shows continue to survive even in today’s internet-connected world. The International Broadcasting Convention returns to the RAI Amsterdam Friday 10 September through Monday 13 September 2027.


Grass Valley Launches AVORA at IBC2026
New standalone solution gives production teams everything they need to create, produce …
Grass Valley EDIUS Future Assurance Program launched at IBC2026
New program gives EDIUS customers access to the next major version at …

Enjoying the news? Sign up for the Creative COW Newsletter!

Sign up for the Creative COW newsletter and get weekly updates on industry news, forum highlights, jobs, inspirational tutorials, tips, burning questions, and more! Receive bulletins from the largest, longest-running community dedicated to supporting professionals working in film, video, and audio.

Enter your email address, and your first and last name below!

Sign up:

* indicates required

Responses

We use anonymous cookies to give you the best experience we can.
Our Privacy policy | GDPR Policy