Modality-First
Connect. Create.

Connect text, images, audio and video into your own AI workflow.

Open-source core · Your models · Self-hostable

Text
Image
Audio
3D
Video
Every result is material you can use again.
One workflow / Four steps

Start with a document.

Bring in a document.

Keep the document on the canvas and reuse it.

Extract the text.

Extract the text for the next step.

Branch in two directions.

Turn the same text into a summary and speech.

Connect an image.

Combine the audio and a character image into a talking video.

Workflow illustration · Requires compatible pluginsScroll to build the workflow

Think in possibilities.
Build with modalities.

Change the model without rebuilding the workflow.

Why TongFlow starts with modalities →
What would you start with?
Start here Transform Continue from here
Text
Develop an idea
Text
Generate an image
Image
Give it a voice
Audio

Capability examples. Running requires compatible plugins.

Operations

Add. Transform. Combine. Split.

Add

Start with text, a file or a link.

Transform

Turn an image into text, video or another image.

Combine

Combine an image and a voice into video.

Split & batch

Split into parts and process each one.

Scroll to follow the connections

Connect the steps your task needs.

01

Make information easier to use

Document Text Summary Audio
02

Take an image further

Image Transform Video / 3D
03

Find the useful parts

Video Clips Audio Text

MODALITY-FIRST

Materials, capabilities and models.

Every result is a new beginning

Connect uploaded and generated materials to the next step.

Choose a capability. Then a model.

Switch compatible models while keeping your workflow.

A shared structure for people and agents

Build on the canvas or let an agent create an editable workflow.

Scroll to follow the connections

One image. One motion reference.

Combine both inputs with a compatible motion-transfer plugin.

Explore the source
TongFlow canvas connecting a cat image and a reference dance video to a motion-transfer operation
Watch the example

Cloud, self-hosted, or built with code.

One open-source core. Connect the plugins and providers you need.

Self-host

Configure plugins and models on your own infrastructure.

Get the source
Get the macOS / Windows app ↗The desktop app opens the cloud workspace; it requires a connection.

A few things to know

What is TongFlow?

TongFlow is an open-source tool for building multimodal AI workflows. Connect text, images, audio, video, 3D and documents on a visual canvas to extract, transform, generate or combine information. Each step runs through a compatible plugin and its model or processing implementation. You can use the cloud workspace or self-host the app.

What does modality-first mean?

You build around forms of information, such as text, images or sound. Capabilities connect those forms; models implement the capabilities. This gives different workflows a common structure.

Can I connect anything to anything?

Connections follow the input and output types of each capability. A workflow also needs compatible plugins. The design is open-ended, but available operations determine what can run today.

Does it only generate media?

No. Understanding an image, extracting text, transcribing speech, splitting a video and combining materials are also part of the system. Generation is one kind of operation.

What do I need to run a workflow?

Use the cloud workspace or self-host the app, then configure the provider credentials required by your chosen plugins. Modal-backed plugins require a Modal account; API plugins require the relevant provider keys.

How does billing work?

Cloud access is a subscription. Model API usage and GPU compute are billed separately by your providers. TongFlow does not sell generation credits. See pricing for the current subscription and trial details.

Pricing →

Connect your first nodes.