---
title: Fed-Star in Astrophysics & Federated Learning
url: https://www.emergentmind.com/topics/fed-star
type: topic
---

# Fed-Star in Astrophysics & Federated Learning

“Fed-Star” and “FedSTAR” denote several distinct constructs in astrophysics and federated learning rather than a single unified concept. Across these usages, the recurring motif is external feeding: cold streams feeding galaxies, stellar winds feeding accretion flows, clump-scale inflow feeding massive protostellar fragments, or exchanged representations feeding federated models. In galaxy formation, the term describes delayed star formation in high-redshift stream-fed galaxies [1310.1923]. In compact-object and stellar contexts, it appears in wind-fed accretion onto the supermassive black hole M31*, clump-fed accretion in high-mass star-forming objects, and wind-fed disks in binaries [2506.04778] [2301.09917] [1903.01873]. In machine learning, FedSTAR names both a semi-supervised federated self-training method for audio recognition and a personalized federated-learning framework based on style-aware prototype aggregation [2107.06877] [2511.18841].

## 1. Nomenclature and scope

| Domain | Meaning of “Fed-Star” / “FedSTAR” | Central mechanism |
|---|---|---|
| Galaxy formation | Delayed star formation in high-redshift stream-fed galaxies | Inflow-driven turbulence suppresses star formation |
| SMBH accretion | Stellar-wind feeding of M31* | AGB-star winds build a cool quasi-Keplerian disk |
| Massive star formation | Clump-fed accretion mechanism | Parsec-scale inflow sustains fragment growth |
| Binary accretion | Wind-fed accretion disk | Red-giant wind feeds a thin disk around a companion |
| Federated learning | FEderated Self-TRAining | Pseudo-labeling exploits on-device unlabeled audio |
| Personalized FL | Federated Style-Aware Transformer Aggregation of Representations | Content–style disentanglement and attention-weighted prototype fusion |

The arXiv record therefore uses the same lexical label for unrelated problems. In astrophysics, the term is attached to feeding mechanisms in gaseous systems; in federated learning, it functions as an acronym. This suggests a mnemonic convergence rather than a standardized cross-disciplinary taxonomy.

A common misconception is to treat “Fed-Star” as a singular model family. The cited literature does not support that reading. Instead, each usage is domain-specific, with independent definitions, observables, and mathematical formalisms.

## 2. Stream-fed suppression of star formation in high-redshift galaxies

In “Delayed star formation in high-redshift stream-fed galaxies,” the Fed-Star mechanism proposes that star formation is delayed relative to the inflow rate in rapidly accreting galaxies at very high redshift because the accreting gas conveys energy into the disk and raises turbulence above the level compatible with gravitational instability [1310.1923]. The inflowing gas therefore acts simultaneously as fuel and as a stabilizing agent.

The analytic model begins from turbulent energy injection by cold streams. For an inflow rate $\dot M_{\rm inflow}$, infall velocity $v_{\rm infall}\simeq \sqrt{2}\,v_{\rm halo}$, and coupling fraction $\epsilon$, the injection rate is
$$
\dot E_{\rm in}=\tfrac12\,\epsilon\,\dot M_{\rm inflow}\,v_{\rm infall}^2,
$$
which yields the scaling
$$
\sigma_{\rm turb}\simeq \bigl[\epsilon\,(\dot M/M_{\rm gas})\bigr]^{1/2}R_{\rm gal}.
$$
Internal processes enforce a floor
$$
\sigma_{\min}=\frac{Q_{\min}\,\pi G\,\Sigma_{\rm gas}}{\kappa},
\qquad Q_{\min}\approx 0.7,
$$
and the actual dispersion is
$$
\sigma=\max(\sigma_{\min},\sigma_{\rm turb}),
$$
so that the instantaneous Toomre parameter becomes
$$
Q=\frac{\kappa\,\sigma}{\pi G\,\Sigma_{\rm gas}}.
$$
Whenever inflow-driven turbulence dominates, $Q$ rises above unity and the disk is stabilized against fragmentation.

The star-formation law is then modified through the density PDF. For a log-normal PDF with width
$$
\sigma_{\ln\rho}^2=\ln\bigl[1+b^2(\sigma/c_s)^2\bigr],
$$
the efficiency per free-fall time is written
$$
\epsilon_{\rm SF}\simeq \epsilon_0\exp\Bigl[-\tfrac32\,\sigma_{\ln\rho}^2\Bigr],
\qquad \epsilon_0\sim 0.01.
$$
This enters a Kennicutt-style law of the form
$$
\Sigma_{\rm SFR}=\epsilon_{\rm SF}\,\frac{\Sigma_{\rm gas}}{t_{\rm orb}}
\quad\text{or}\quad
\Sigma_{\rm SFR}=A\,\Sigma_{\rm gas}^N\,\sigma^{-1},
$$
with $N\approx 1$–$1.5$ when turbulence suppresses collapse. The gas fraction is
$$
f_g=\frac{M_{\rm gas}}{M_{\rm gas}+M_*}.
$$

The redshift dependence is central. At $z>2$, theoretical accretion rates scale as $(1+z)^{2.25}$, so low-mass galaxies experience very high $(\dot M/M)_{\rm inflow}\sim {\rm few}\ {\rm Gyr}^{-1}$. For $R_{\rm gal}\sim 2\,{\rm kpc}$ and $\epsilon\sim 0.5$–$1$, the model gives $\sigma_{\rm turb}\gg \sigma_{\min}$ and hence $Q\gtrsim 1.5$–$2.0$. The star-formation efficiency is reduced by a factor of about three relative to the self-regulated floor, and $f_g\gtrsim 0.4$–$0.6$ is maintained down to $z\sim 2$. As $z$ drops below $\sim 2$–$3$, the specific inflow rate falls by $\gtrsim \times 3$ and the geometric coupling factor $A_{\rm infall}$ decreases as filaments decouple from the compact disk. Then $\sigma\rightarrow \sigma_{\min}$, $Q\rightarrow Q_{\min}\approx 0.7$–$1$, $\epsilon_{\rm SF}$ returns to its canonical $\sim 0.01$, and $f_g$ declines toward $10$–$20\%$ by $z\sim 1$.

Idealized hydrodynamic simulations with RAMSES at $32\,{\rm pc}$ resolution down to $n\approx 0.2\,{\rm cm}^{-3}$ support the analytic picture. At $z\approx 5$, runs with $M_*\sim 10^9\,M_\odot$, $f_g=0.75$, and three filamentary streams totaling $\dot M=15\,M_\odot\,{\rm yr}^{-1}$ yield a coupling efficiency $\epsilon\approx 0.8$–$1.0$. At $z\approx 2$, analogous runs with $\dot M\simeq 80\,M_\odot\,{\rm yr}^{-1}$, $M_*=8\times 10^{10}\,M_\odot$, and $f_g=0.5$ give $\epsilon\lesssim 0.2$. In the $z>2$ simulations, the stream-fed case has $f_g\approx 0.6$ and ${\rm SFR}\approx 10\,M_\odot\,{\rm yr}^{-1}$, compared with $f_g\approx 0.35$ and ${\rm SFR}\approx 25\,M_\odot\,{\rm yr}^{-1}$ in the control. The global efficiency ${\rm SFR}/M_{\rm gas}$ drops from $\simeq 0.015\,{\rm Gyr}^{-1}$ in the control to $0.005\,{\rm Gyr}^{-1}$ in the fed run, while $\Sigma_{\rm SFR,stream}\simeq 0.3\,\Sigma_{\rm SFR,control}$ at fixed $\Sigma_{\rm gas}$.

Relative to traditional bathtub or self-regulated models, this framework predicts a prolonged gas-rich phase, suppressed early stellar-mass build-up, thicker high-$\sigma$ disks with $\sigma\sim 70$–$100\,{\rm km\,s}^{-1}$, and a transition near $z\approx 2$–$3$ to marginally stable star formation. The paper explicitly frames this as a way to unify high gas fractions, elevated dispersions, delayed star formation, and the later self-regulated regime within one stream-feeding picture.

## 3. Stellar-wind feeding of M31*

For M31*, Fed-Star denotes a stellar-wind feeding mechanism in which the central supermassive black hole is supplied by collective mass loss from the surrounding nuclear star cluster [2506.04778]. The mass-losing population is modeled as $\sim 100$ thermally-pulsing AGB stars associated with an $\sim 8\,{\rm Gyr}$-old, metal-rich population with $[{\rm M/H}]\approx +0.3$ and total mass $M_{\rm NSC}\approx 2\times 10^7\,M_\odot$.

Each AGB star is assigned a time-averaged mass-loss rate $\dot M_\star\approx 4\times 10^{-7}\,M_\odot\,{\rm yr}^{-1}$, wind temperature $T_w\approx 3000\,{\rm K}$, and wind speed $v_w\approx 10\,{\rm km\,s}^{-1}$. The ensemble therefore injects $\dot M_{\rm inj}\approx 4\times 10^{-5}\,M_\odot\,{\rm yr}^{-1}$. The stars move on Keplerian orbits around a central SMBH of mass $M_\bullet=1\times 10^8\,M_\odot$, sampling orbital elements with $a\simeq 3.3\,{\rm pc}$, $e\simeq 0.35$, and inclination $\simeq 54^\circ$. Winds are injected within a sphere of radius $r_{\rm inj}\approx 0.16\,{\rm pc}$ centered on each star and carry both orbital velocity, about $400\,{\rm km\,s}^{-1}$, and intrinsic wind velocity.

The simulations solve the Euler equations with source terms for wind mass, momentum, and energy injection, together with external gravity and radiative heating/cooling:
$$
\frac{\partial \rho}{\partial t}+\nabla\cdot(\rho v)=\dot\rho_w,
$$
$$
\frac{\partial(\rho v)}{\partial t}+\nabla\cdot(\rho v v+pI)=-\rho\nabla\Phi+\dot m_w,
$$
$$
\frac{\partial E_t}{\partial t}+\nabla\cdot[(E_t+p)v]=-\rho v\cdot\nabla\Phi+\dot\rho_w\Phi+\dot E_w+\dot Q.
$$
Here $E_t=p/(\gamma-1)+\tfrac12\rho v^2$ with $\gamma=5/3$, and $\dot Q$ is derived from CLOUDY-based lookup tables.

The numerical setup uses PLUTO 4.4 on a Cartesian grid of $16\times 16\times 8\,{\rm pc}^3$ with $384\times 384\times 192$ zones, corresponding to $\Delta\approx 0.04\,{\rm pc}$. The innermost $2^3$ cells define an effective accretion radius $r_{\rm acc}\approx 8700\,r_g$. The fiducial and point-mass runs use outflow boundaries, whereas the inflow run adds an isotropic inflow of $\dot M_{\rm in}\approx 1.6\times 10^{-5}\,M_\odot\,{\rm yr}^{-1}$, $T_{\rm in}=10^5\,{\rm K}$, and $v_{\rm in}=100\,{\rm km\,s}^{-1}$. The evolution is followed for $2.0\,{\rm Myr}$ with a time step of about $1\,{\rm yr}$.

By $t\approx 2\,{\rm Myr}$, the slow and cold AGB winds have collided, shock-heated, radiatively cooled, and settled into a flattened eccentric disk in the mean orbital plane. The disk extends to $r\approx 3$–$6\,{\rm pc}$, with $T_{\rm disk}\approx 10^3$–$10^4\,{\rm K}$ and $n_{\rm disk}\sim 10^2$–$10^4\,{\rm cm}^{-3}$. It is embedded in a hot halo with $n_{\rm halo}\sim 10^{-4}$–$10^{-3}\,{\rm cm}^{-3}$ and $T_{\rm halo}\sim 10^6$–$10^7\,{\rm K}$. The surface density declines roughly as $\Sigma(r)\propto r^{(-1\pm 0.5)}$ and peaks near $\Sigma\sim 10^3\,M_\odot\,{\rm pc}^{-2}$ at small radii. The scale height obeys
$$
H(r)\simeq \frac{c_s}{\Omega_K},
$$
giving $H/r\sim 0.05$–$0.2$ for $T\sim 10^4\,{\rm K}$ across $1$–$5\,{\rm pc}$.

The accretion rate through $r_{\rm acc}$ approaches a quasi-steady value. The point-mass run yields $\dot M_{\rm acc}\approx 6.8\times 10^{-6}\,M_\odot\,{\rm yr}^{-1}$, or $17\%$ of $\dot M_{\rm inj}$; the fiducial run gives $\dot M_{\rm acc}\approx 2.4\times 10^{-5}\,M_\odot\,{\rm yr}^{-1}$, or $60\%$ of $\dot M_{\rm inj}$; and the inflow run reaches $\dot M_{\rm acc}\approx 5.4\times 10^{-5}\,M_\odot\,{\rm yr}^{-1}$, or $135\%$ of $\dot M_{\rm inj}$. The non-axisymmetric NSC potential increases $\dot M_{\rm acc}$ by about $3.5\times$, and short-term fluctuations of order $50\%$ track stars passing pericenter on $\sim 10^4\,{\rm yr}$ timescales.

The predicted observables include an X-ray luminosity $L_X\approx (0.8$–$3)\times 10^{36}\,{\rm erg\,s}^{-1}$ from hot plasma within $\sim 0.2\,{\rm pc}$, consistent with the Chandra range $L_X\approx 2\times 10^{35}$–$4\times 10^{36}\,{\rm erg\,s}^{-1}$. The synthetic spectrum would appear very soft if fitted by a power law, with photon index $\Gamma\sim 4.2$. Photoionization of the cool disk gives $L_{{\rm H}\alpha}\approx 1.4\times 10^{36}\,{\rm erg\,s}^{-1}$, comparable to the observed $\sim (3.4\pm 0.4)\times 10^{36}\,{\rm erg\,s}^{-1}$, and predicts optical forbidden lines and IR lines potentially accessible to JWST. The paper concludes that old-star winds can dominate SMBH fueling in quiescent nuclei and argues that cosmological and galaxy-evolution simulations should include NSC wind feeding as a sub-grid source term.

## 4. Clump-fed accretion in high-mass star-forming objects

Within the SQUALO project, Fed-Star refers to a clump-fed mechanism for the formation of massive stars [2301.09917]. The observational basis is an ALMA Band 6 and Band 3 continuum survey of 13 massive clumps selected from the Hi-GAL and MALT90 catalogues for having blue-asymmetric HCO$^+$(1–0) or HNC(1–0) profiles indicating infall. The selection requires $M_{\rm parent}\ge 170\,M_\odot$, $\Sigma_{\rm parent}\ge 1\,{\rm g\,cm}^{-2}$, $d\le 5.5\,{\rm kpc}$, and relative isolation. Three additional $70\,\mu{\rm m}$-quiet clumps with $L/M<1$ and infall signatures extend the sample over $L/M\approx 0.1$–$107$.

The ALMA data combine 12 m and 7 m arrays in single-pointing mosaics, with typical synthesized beam $1.0''$–$1.3''$, corresponding to $0.01$–$0.04\,{\rm pc}$ or $2000$–$8000\,{\rm AU}$, and rms noise $0.8$–$15\,{\rm mJy\,beam}^{-1}$. All clumps have single-dish infall rates $\dot M_{\rm par}\simeq 0.7$–$27\times 10^{-3}\,M_\odot\,{\rm yr}^{-1}$.

The fragment mass is derived from the $1.3\,{\rm mm}$ continuum via
$$
M_f=\frac{D^2\,S_{1.3}}{\kappa_{1.3}\,B_{1.3}(T_f)},
$$
with $\kappa_{1.3}=0.005\,{\rm cm}^2\,{\rm g}^{-1}$. Surface density is
$$
\Sigma=\frac{M}{A},\qquad A=\pi R^2.
$$
Thermal Jeans scales are written as
$$
\lambda_J=c_s\sqrt{\frac{\pi}{G\rho}},
\qquad
c_s=\sqrt{\frac{k_B T_{\rm cl}}{\mu m_H}},
\qquad
\rho=\frac{M_c}{\frac43\pi R_c^3},
$$
and
$$
M_J=\frac{4\pi}{3}\rho\Bigl(\frac{\lambda_J}{2}\Bigr)^3.
$$
The clump-formation efficiency is
$$
{\rm CFE}=\frac{\sum_i M_{f,i}}{M_c},
$$
and the virial parameter is
$$
\alpha_{\rm vir}=\frac{5\,\sigma_v^2\,R_c}{G\,M_c}.
$$

The survey identifies 55 fragments in 13 clumps, with $0.4\lesssim M_f\lesssim 309\,M_\odot$. All three $70\,\mu{\rm m}$-quiet clumps already contain $2$–$4$ fragments, which the authors interpret as evidence that massive “starless” cores are rare. One source, HIGALBM343.7560–0.1629, with $L/M\approx 14.7$, hosts a single $M_f\approx 73\,M_\odot$ object. The fragment and clump properties are correlated: $M_{f,\max}\propto M_c$ with $\rho\approx 0.73$, $\Sigma_{f,\max}\propto \Sigma_c$ with $\rho\approx 0.60$, total fragment mass correlates weakly with $\dot M_c$ with $\rho\approx 0.44$, and $\alpha_{\rm vir,c}\propto M_f^{-0.27}$ with $\rho\approx -0.41$.

Fragment spacing evolves systematically. The minimum projected separation $d_{\min}$ decreases as $L/M$ increases: in early clumps $d_{\min}\gtrsim 0.03\,{\rm pc}$, whereas in evolved systems fragments reach separations of order $1000\,{\rm AU}$. Jeans analysis shows that the thermal Jeans ratio $\lambda_{J,r}\equiv d_{\min}/\lambda_{J,{\rm thermal}}\gg 1$ in young clumps and approaches unity in evolved clumps, while the non-thermal Jeans ratio $\lambda_{J,r,{\rm nth}}<1$ in almost all clumps. The observational interpretation is therefore staged. Early fragmentation is “gravo-turbulent,” with large-scale turbulence and gravity producing a small number of massive fragments at scales larger than the thermal Jeans length. As collapse proceeds, turbulence dissipates or infall accelerates, separations shrink, and fragmentation approaches the thermal Jeans scale. Magnetic support is invoked for the non-fragmenting source as a special case.

The proposed clump-fed scenario has five steps: parsec-scale gas inflow with $\dot M_c\gtrsim 10^{-3}\,M_\odot\,{\rm yr}^{-1}$ drives global collapse; turbulence seeds a handful of massive fragments at $\sim 0.05$–$0.1\,{\rm pc}$; continuous accretion from the clump raises fragment mass and surface density; over $\sim 10^5\,{\rm yr}$ turbulence is damped and fragments contract to separations of $\sim 10^3\,{\rm AU}$; embedded protostars then continue to accrete from the common clump reservoir along filaments. The paper explicitly contrasts this hierarchical, multi-scale accretion picture with a pure core-fed model.

## 5. Wind-fed accretion disks and planet migration in binaries

In binary-star accretion, Fed-Star denotes a wind-fed disk formed when a secondary captures part of the slow dense wind of a red-giant companion through Bondi–Hoyle accretion [1903.01873]. The analysis assumes that the disk viscous time is shorter than the wind-variation time, allowing a quasi-steady $\alpha$-disk treatment.

The disk is geometrically thin and Keplerian, with
$$
\Omega(r)=\sqrt{\frac{G M_2}{r^3}},
$$
scale height
$$
H(r)=\frac{c_s}{\Omega},
$$
and viscosity
$$
\nu(r)=\alpha c_s H.
$$
Two feeding geometries are considered. In the standard disk, matter is supplied at the outer edge and the accretion rate is radially constant:
$$
\dot M_{\rm acc}(r)=\dot M_{\rm acc}^{\rm tot}.
$$
Angular-momentum conservation gives
$$
\nu\Sigma=\frac{\dot M_{\rm acc}^{\rm tot}}{3\pi}f(r),
\qquad
f(r)=1-\sqrt{\frac{r_{\rm in}}{r}},
$$
and radiative balance yields
$$
T_c(r)^4=\frac{27}{64\sigma}\,\kappa\,\nu\,\Omega^2\,\Sigma^2.
$$
Far from $r_{\rm in}$, the standard scalings are
$$
\Sigma_{\rm SD}(r)\propto \alpha^{-4/5}M_2^{1/5}\dot M_{\rm acc}^{3/5}r^{-3/5},
$$
$$
T_{\rm SD}(r)\propto \alpha^{-1/5}M_2^{3/10}\dot M_{\rm acc}^{2/5}r^{-9/10}.
$$

In the distributed wind-fed case, material settles over all radii $r<r_a$ at a rate
$$
\dot\Sigma_{\rm ext}(r)=
\begin{cases}
\dfrac{\dot M_{\rm acc}^{\rm tot}}{2\pi r r_a}, & r<r_a,\\
0, & r\ge r_a,
\end{cases}
$$
so that the same $\nu\Sigma$ relation holds but with modified $f(r)$. In the regime $r_{\rm in}\ll r\ll r_a$,
$$
f(r)\simeq 1+\frac{r}{3r_a}.
$$
The only formal difference from the standard solution is therefore an extra factor $f(r)^{3/5}$ in $\Sigma$ and $f(r)^{2/5}$ in $T$.

Planet migration is treated in the classical Type I/II framework. For Type I migration in a three-dimensional isothermal disk with $\Sigma\propto r^{-\bar\alpha}$, the torque is
$$
\Gamma_{\rm I}
=
-\bigl[1.36+0.54\,\bar\alpha\bigr]
\Bigl(\frac{M_p}{M_2}\Bigr)^2
\Bigl(\frac{r_p}{H}\Bigr)^2
\Sigma(r_p)\,r_p^4\,\Omega_p^2,
$$
which implies
$$
\frac{dr_p}{dt}
=
-\bigl[2.72+1.08\,\bar\alpha\bigr]
\Bigl(\frac{M_p}{M_2}\Bigr)^2
\Bigl(\frac{r_p}{H}\Bigr)^2
\Sigma(r_p)\,r_p^3\,\Omega_p,
$$
and $\tau_{\rm mig,I}=|r_p/(dr_p/dt)|$. Gap opening and Type II migration are described by the criterion
$$
q\equiv \frac{M_p}{M_2}>q_{\rm crit},
$$
with $q_{\rm crit}$ expressed in terms of the Reynolds number $\mathcal R=r_p^2\Omega_p/\nu$, after which the drift rate is
$$
\frac{dr_p}{dt}
=
-\frac32\,\alpha\Bigl(\frac{H}{r_p}\Bigr)^2 f^{-1}(r_p)v_K,
$$
and
$$
\tau_{\rm mig,II}\simeq \frac23\,\frac{r_p^2}{\nu}.
$$

For red-giant mass-loss rates $\dot M_w\sim 10^{-9}$–$10^{-6}\,M_\odot\,{\rm yr}^{-1}$ and binary separations $a\sim 10$–$50\,{\rm AU}$, the capture rate is $\dot M_{\rm acc}\sim 10^{-11}$–$10^{-8}\,M_\odot\,{\rm yr}^{-1}$. With $\alpha\sim 10^{-3}$–$10^{-2}$, the disk lifetime is set by the red-giant phase, $t_{\rm disk}\sim 10^7\,{\rm yr}$. In standard edge-fed disks, Type I migration at $1\,{\rm AU}$ is $\tau_{\rm mig,I}\sim 10^6$–$10^8\,{\rm yr}$ for $M_p\sim 0.01$–$1\,M_{\rm J}$ if $\dot M_{\rm acc}\gtrsim 10^{-9}\,M_\odot\,{\rm yr}^{-1}$; for lower accretion rates the disk is too tenuous for migration within the disk lifetime. The Type I–Type II transition occurs at $M_p\sim 0.1$–$1\,M_{\rm J}$, and Type II migration is generally faster in these low-mass disks. Jupiter-mass planets can merge within $10^6$–$10^7\,{\rm yr}$ for $a\lesssim 30\,{\rm AU}$ and $\dot M_w\gtrsim 10^{-7}\,M_\odot\,{\rm yr}^{-1}$, whereas lower-mass planets may survive if $\dot M_w\lesssim 10^{-8}\,M_\odot\,{\rm yr}^{-1}$ or $a\gtrsim 50\,{\rm AU}$.

The disk surface densities are much lower than in protoplanetary disks, about $\Sigma\sim 1$–$10\,{\rm g\,cm}^{-2}$ at $1\,{\rm AU}$ versus $\Sigma\sim 10^2$–$10^3\,{\rm g\,cm}^{-2}$, which slows Type I migration and raises the critical mass for gap opening. Yet the longer wind-fed disk lifetime means that substantial migration remains possible. The merger energy $\Delta E\sim G M_2 M_p/R_2\sim 10^{44}$–$10^{46}\,{\rm erg}$ motivates the transient interpretation discussed in the paper.

## 6. FEderated Self-TRAining for semi-supervised audio recognition

In machine learning, FedSTAR was introduced as “FEderated Self-TRAining” for semi-supervised audio recognition [2107.06877]. The method addresses federated learning with scarce labeled audio and abundant unlabeled audio distributed across devices. The goal is to train a single global model while keeping raw audio local and exploiting pseudo-labeling on each client.

The per-round workflow is straightforward. The server maintains global parameters $\theta_r^G$, samples a fraction $q$ of clients, and sends $\theta_r^G$ to each selected client. Client $k$ performs $E$ local epochs using labeled minibatches from $\mathcal D_k^L$ and unlabeled minibatches from $\mathcal D_k^U$. The local objective combines supervised cross-entropy with pseudo-label-based unsupervised cross-entropy, where low-confidence pseudo-labels are discarded through a dynamic threshold $\tau_r$. Updated local models are then aggregated by weighted FedAvg:
$$
\theta_{r+1}^G \leftarrow \sum_k \frac{N_k}{N}\,\theta_r^k.
$$

The global optimization problem is
$$
\min_\theta L(\theta)=\sum_{k=1}^K \gamma_k L_k(\theta),
\qquad
\gamma_k=\frac{N_k}{N},
$$
with local loss
$$
L_k(\theta)=L_s(\theta;\mathcal D_k^L)+\beta\,L_u(\theta;\mathcal D_k^U).
$$
The supervised term is categorical cross-entropy,
$$
L_s(\theta;\mathcal D_k^L)=
-\frac{1}{N_{\ell,k}}
\sum_{i\in \mathcal D_k^L}\sum_{j=1}^C y_{ij}\log p_\theta(j\mid x_i^\ell),
$$
while the pseudo-label is obtained from temperature-scaled logits
$$
\hat y(x^u)=\Phi(z,T)=\arg\max_i \left\{\frac{e^{z_i/T}}{\sum_j e^{z_j/T}}\right\},
$$
and the unsupervised term is
$$
L_u(\theta;\mathcal D_k^U)=
-\frac{1}{N_{u,k}}
\sum_{i\in \mathcal D_k^U}\sum_j \hat y_{i,j}\log p_\theta(j\mid x_i^u).
$$

The framework optionally initializes the model with a self-supervised encoder trained on a large unlabeled corpus such as FSD-50K using an InfoNCE-style objective on paired segments from the same clip:
$$
L_{\rm ssl}=
-\sum_i
\log
\frac{\exp({\rm sim}(z_i,z_i')/\tau_c)}
{\sum_j \exp({\rm sim}(z_i,z_j')/\tau_c)}.
$$
This pretrained encoder becomes $\theta_0^G$ and is reported to reduce the number of required federated rounds by $3\times$–$5\times$ for the same accuracy.

Experiments use Ambient Acoustic Context, Speech Commands v2, and VoxForge, with audio resampled to $16\,{\rm kHz}$ and represented as $1\,{\rm s}$ log-Mel spectrograms with $64$ Mel bins. The model has four convolutional blocks, each comprising a timewise $1$D convolution, a frequencywise $1$D convolution, concatenation, a $1\times 1$ convolution, GroupNorm, ReLU, $L_2$ weight decay $=10^{-4}$, spatial dropout $=0.1$, and max-pooling $2\times 2$ between blocks, followed by global average pooling and a dense softmax head. Training uses Adam with $\eta=10^{-3}$ and client batch size about $32$. The federation parameters span $N\in \{5,10,15,30\}$ clients, $q=20$–$80\%$, $E=1$–$4$, labeled fraction $L\in\{3\%,5\%,20\%,50\%,100\%\}$, unlabeled fraction $U\in\{20\%,50\%,80\%,100\%\}$, $\beta=0.5$, $T=4$, and a cosine-rising threshold $\tau_r$ from $0.5$ to $0.9$.

Quantitatively, with only $L=3\%$ labels and $N=15$, performance improves from $46.3\%$ to $49.5\%$ on Ambient Context, from $62.98\%$ to $86.82\%$ on Speech Commands, and from $54.3\%$ to $55.8\%$ on VoxForge. Averaged across tasks and client counts, the method improves recognition by up to $13.28\%$ over fully supervised federated learning at $L=3\%$. Under extreme non-IIDness, where each client sees only $3/12$ classes on $\mathcal D_k^L$, supervised federated learning remains below $25\%$, whereas FedSTAR still reaches about $80$–$85\%$ for $L=3$–$50\%$. After $10$ federated rounds with $N=15$ and $L=50\%$, SSL initialization improves Speech Commands from about $92\%$ to about $95\%$.

The paper characterizes the method as a lightweight extension of FedAvg because clients need only add a pseudo-label cross-entropy term with tunable $\beta$, $T$, and $\tau$. The principal claim is not personalization but better use of on-device unlabeled data under label scarcity.

## 7. Style-aware transformer aggregation in personalized federated learning

A distinct 2025 usage, “Federated Style-Aware Transformer Aggregation of Representations,” also abbreviated FedSTAR, targets personalized federated learning under domain heterogeneity, data imbalance, and communication constraints [2511.18841]. The central claim is that client embeddings entangle task-relevant content with client-specific style and that uniform averaging of class-wise prototypes suppresses minority-client signals.

Each client extracts features $h_i(x)$ through a shared encoder and maintains, for every class $c$, a mean feature prototype
$$
m_{k,c}=\frac{1}{|\mathcal D_{k,c}|}\sum_{x\in \mathcal D_{k,c}} h_k(x)\in \mathbb R^d
$$
together with a personal residual parameter $u_{k,c}\in \mathbb R^d$. Relative to the current global prototype $p_c^{\rm global}$, the residual is decomposed into content and style. The content projection is
$$
p^{\rm content}_{k,c}
=
\frac{u_{k,c}^\top p_c^{\rm global}}
{\|p_c^{\rm global}\|^2+\varepsilon}\,p_c^{\rm global},
$$
while the orthogonal style residual is
$$
r_{k,c}=u_{k,c}-p^{\rm content}_{k,c},
\qquad
s_{k,c}=\frac{r_{k,c}}{\|r_{k,c}\|+\varepsilon}.
$$
The full local prototype is
$$
p_{k,c}=m_{k,c}+u_{k,c}.
$$
For communication, clients send only the content portion $m_{k,c}+p^{\rm content}_{k,c}$, or equivalently just $p^{\rm content}_{k,c}$ when $m_{k,c}$ is shared. The style vectors remain local and are used for FiLM-based personalization during inference.

On the server, class-wise content prototypes from $M$ participating clients are stacked into a tensor $CP\in \mathbb R^{M\times C\times d}$. Tokens are formed as
$$
X_{k,c}=
{\rm LayerNorm}\bigl(
p^{\rm shared}_{k,c}+e_k^{\rm client}+e_c^{\rm class}
\bigr),
$$
with learned client and class embeddings. A standard Transformer encoder is then applied:
$$
Q=XW^Q,\qquad K=XW^K,\qquad V=XW^V,
$$
$$
Z={\rm Softmax}\!\Bigl(\frac{QK^\top}{\sqrt d}\Bigr)V.
$$
A second class-driven attention computes
$$
\alpha_{k,c}
=
\frac{\exp(Z_{k,c}^\top e_c^{\rm class}/\sqrt d)}
{\sum_{j=1}^M \exp(Z_{j,c}^\top e_c^{\rm class}/\sqrt d)},
$$
and the updated global prototype is
$$
p_c^{\rm global}=\sum_{k=1}^M \alpha_{k,c} Z_{k,c}.
$$
Clients then fuse the global prototypes with local residual parameters through a learned gating network.

Communication efficiency is a primary design goal. Rather than exchanging full model weights of size $O(P)$, each client sends $C$ content prototypes of dimension $d$, and optionally $C$ style vectors, for total communication $O(2Cd)$. The ratio
$$
r=\frac{2Cd}{P}\ll 1
$$
is reported as typically $r\le 5\%$, amounting to one to two orders of magnitude less communication than full-model exchange in typical settings.

The evaluation uses Fashion-MNIST, CIFAR-100, DomainNet, and Office-31 under severe non-IID Dirichlet splits with $\alpha=0.1$ plus Gaussian noise. The reported results are: on Fashion-MNIST, FedProto achieves $85.80\%\pm 0.18$ accuracy, the attention-only ablation reaches $86.57\%$, and FedSTAR reaches $89.37\%\pm 0.18$ with $F1=0.8936$ and convergence in $55$ rounds; on CIFAR-100, performance increases from $24.41\%\pm 0.09$ for FedProto to $26.29\%$ for the ablation and $29.17\%$ for FedSTAR; on DomainNet, from $14.52\%$ to $14.91\%$ to $16.03\%$; and on Office-31, from $41.92\%$ to $43.65\%$ to $45.03\%$. Ablations attribute a $1$–$2$ percentage-point gain to replacing uniform averaging with Transformer attention alone and a further $2$–$3$ percentage-point gain to adding style-aware FiLM personalization.

This framework differs sharply from the audio self-training FedSTAR despite the identical acronym. One addresses semi-supervised learning with pseudo-labels and a single global model; the other addresses personalized federated learning through explicit content–style disentanglement and attention-weighted prototype aggregation. The shared acronym does not indicate methodological continuity.

Source: https://www.emergentmind.com/topics/fed-star