Finite-Horizon Joint Power Allocation (FHJPA)
- Finite-Horizon Joint Power Allocation (FHJPA) is a time-limited control strategy that jointly optimizes source transmit and destination jamming powers over known decision epochs.
- The method employs backward induction and a finite Markov decision process to manage battery causality, channel uncertainties, and energy-harvesting constraints.
- Studies indicate FHJPA achieves superior secrecy energy efficiency compared to greedy or infinite-horizon approaches, with structural insights like threshold limits and water-filling behaviors.
Finite-Horizon Joint Power Allocation (FHJPA) denotes a class of stage-dependent wireless control problems in which power variables are optimized jointly over a known and finite set of decision epochs, rather than by a stationary infinite-horizon rule or by purely myopic slot-wise maximization. In the specific sense used by "Secure Energy Efficient Wireless Transmission: A Finite v/s Infinite-Horizon RL Solution" (Tripathi et al., 4 Aug 2025), FHJPA is the backward-induction policy that jointly allocates the source transmit power and the destination jamming power in an energy-harvesting wiretap system so as to maximize average secrecy energy efficiency over a fixed transmission horizon. Related finite-horizon formulations in the literature include joint optimization of wireless power transfer duration and uplink power allocation (Abad et al., 2019), hard-deadline scheduling and power control in downlink cellular systems (Ewaisha et al., 2017), and adaptive distributed power allocation with terminal constraints under channel uncertainty in cognitive radio networks (Xu et al., 2013). Taken together, these works place FHJPA within the broader theory of finite-horizon stochastic control for wireless systems with causality, deadline, interference, and battery constraints.
1. Finite-horizon formulation and defining characteristics
The defining feature of FHJPA is that the number of decision epochs is known in advance and finite. In the secrecy-energy-efficiency formulation, communication occurs over slots
each of duration , and the optimal rule is explicitly stage dependent because the value of preserving energy changes as the terminal slot approaches (Tripathi et al., 4 Aug 2025). The same finite-horizon logic appears in the wirelessly powered-device formulation, where the horizon is and the controller must decide both when to stop harvesting and how to allocate battery energy over the remaining uplink slots (Abad et al., 2019).
A finite-horizon model is used when the control problem is inherently time-limited. In the secrecy setting, the finite-horizon formulation is described as most appropriate when the number of decision epochs is known and the system is non-stationary across stages; by contrast, infinite-horizon reinforcement learning assumes an unending or unknown lifetime and yields stationary policies that can be inaccurate when is small or fixed (Tripathi et al., 4 Aug 2025). In the hard-deadline cellular setting, the long-term objective is coupled to a per-slot service horizon because each real-time packet arriving in slot must be fully transmitted within that same slot or is dropped, making the slot-level decision effectively a finite-horizon deadline-constrained control problem (Ewaisha et al., 2017).
The “joint” aspect is not unique to any single system architecture. Depending on the formulation, it can refer to jointly selecting source and jammer powers (Tripathi et al., 4 Aug 2025), jointly optimizing energy-harvesting stopping time and uplink power allocation (Abad et al., 2019), jointly handling scheduling, power allocation, and admission control under deadline constraints (Ewaisha et al., 2017), or jointly allocating power across primary and secondary users while regulating interference and SIR tracking over a finite horizon (Xu et al., 2013).
2. Markov decision structure in the canonical FHJPA algorithm
In the secrecy-energy-efficiency setting, FHJPA is formulated as a finite Markov decision process with finite state and action spaces (Tripathi et al., 4 Aug 2025). The state at slot is
where the channel gains are quantized into finite channel-state sets and the battery levels are quantized into finite battery-state sets. The action is the power pair
selected from a finite power set, so the action space contains possible actions.
The feasible-action constraint is battery-causal:
State transitions factor into Markov channel transitions and battery transitions driven by harvested-energy arrivals. Because the channel power gains for 0, 1, 2, and 3 evolve as a first-order Markov process and the energy arrivals are Bernoulli, the overall system admits an explicit transition kernel (Tripathi et al., 4 Aug 2025).
FHJPA computes the optimal policy by backward induction. At the terminal slot,
4
and for earlier slots,
5
The optimal decision rule 6 is the maximizer of the same Bellman expression. The policy is computed offline over the entire horizon, stored in a lookup table, and then used online by observing the current state, reading off the optimal action, transmitting and jamming accordingly, and updating batteries through the harvesting dynamics (Tripathi et al., 4 Aug 2025).
This stage-indexed Bellman recursion is the core distinction between FHJPA and stationary infinite-horizon methods. The value function depends on the remaining lifetime, so the same physical state can induce different power actions at different stages.
3. Objective, secrecy model, and energy-harvesting constraints
In the canonical FHJPA paper, the network contains an energy-harvesting source 7, an energy-harvesting full-duplex destination 8, and a passive eavesdropper 9 (Tripathi et al., 4 Aug 2025). The source always has data to transmit. The destination simultaneously receives the source signal and emits artificial jamming, which improves secrecy but creates residual self-interference. The cancellation factor satisfies 0, with 1 corresponding to perfect cancellation and 2 to no cancellation.
The secrecy rate is defined per slot as
3
with
4
The destination and eavesdropper SINRs are
5
These expressions expose the central coupling: increasing 6 degrades the eavesdropper but also harms the legitimate receiver through residual self-interference.
The optimization target is the average secrecy energy efficiency,
7
and the per-slot reward is the secrecy energy-efficiency ratio. The batteries have finite capacities 8 and 9, and harvested-energy arrivals are Bernoulli:
0
1
Battery evolution is clipped by capacity, so excess harvested energy is discarded when the maximum is reached (Tripathi et al., 4 Aug 2025).
An important interpretive point follows directly from the stated objective: FHJPA is designed to maximize secrecy energy efficiency, not raw secure throughput. The paper therefore also reports the expected total transmitted secure bits,
2
and notes that FHJPA is not always the best under that metric because higher-power policies can transmit more secure bits while reducing energy efficiency (Tripathi et al., 4 Aug 2025).
4. Structural results in related finite-horizon power-allocation literature
Finite-horizon joint power allocation is not restricted to secrecy problems. The literature exhibits several recurring structures: threshold stopping, closed-form per-slot power laws, deadline-induced inversions of standard water-filling behavior, and adaptive dynamic programming under model uncertainty.
| Work | Joint decision variables | Structural result |
|---|---|---|
| (Abad et al., 2019) | EH stopping time 3 and uplink power allocation 4 | Threshold rule for stopping EH; closed-form 5 |
| (Ewaisha et al., 2017) | Scheduling, power allocation, admission control | Lambert-6 power for RT users; water-filling-like power for NRT users |
| (Xu et al., 2013) | PU/SU power allocation over a finite horizon | Distributed ADP with SIR tracking and UUB error bounds |
In the wirelessly powered-device model, the horizon is partitioned into an energy-harvesting period and an information-transmission period. The controller decides when to stop harvesting and begin transmitting, then allocates a fraction 7 of the current battery energy during each uplink slot (Abad et al., 2019). The dynamic program yields two structural results: all remaining energy is used in the final slot, 8, and the optimal harvesting stop rule is threshold based, with
9
A key lemma establishes that the future-value term 0 is monotonically decreasing in 1, which underpins the threshold structure. This implies that the battery threshold is higher earlier in the horizon and lower near the deadline.
In the hard-deadline downlink problem, the power law depends sharply on traffic class (Ewaisha et al., 2017). For non-real-time users, the optimal power is water-filling-like:
2
which is increasing in channel quality. For real-time users, the power law has a Lambert-3 form and is monotonically decreasing in channel gain. The paper attributes this inversion to the hard-deadline service model: the controller is not opportunistically maximizing bits over time, but rather delivering a fixed packet within the slot under the drift objective.
In the enhanced cognitive-radio formulation, the finite horizon is 4 and the objective is SIR tracking with power minimization under channel uncertainty (Xu et al., 2013). The controller is distributed and model-free, the value function is parameterized linearly in unknown quantities, and the parameter update is driven by Bellman equation residual error and terminal constraint estimation error. The resulting closed-loop guarantees are stated in terms of uniformly ultimately bounded SIR errors.
5. Performance, complexity, and comparison with alternative policies
The FHJPA paper compares the finite-horizon policy against a low-complexity greedy algorithm (GA) and an infinite-horizon joint power allocation method (IHJPA) (Tripathi et al., 4 Aug 2025). GA maximizes only the immediate reward in each slot,
5
whereas IHJPA uses an infinite-horizon discounted formulation and stationary policy iteration. For fairness, the mean effective horizon of IHJPA is taken as
6
The reported computational complexities are:
- FHJPA planning phase: 7
- FHJPA transmission phase: 8
- GA transmission phase: 9
- IHJPA planning phase: 0
- IHJPA transmission phase: 1
The paper explicitly reports that FHJPA takes 2 percent less computational time than IHJPA (Tripathi et al., 4 Aug 2025). In the tested settings, FHJPA achieves the highest average secrecy energy efficiency. GA can approach FHJPA when the source battery has sufficient energy, and the reported trend is that increasing 3 or the harvesting probability 4 brings GA closer to FHJPA. The performance gap between FHJPA and IHJPA decreases as 5 grows, because the infinite-horizon approximation becomes more accurate for longer operational lifetimes.
The related finite-horizon literature reports complementary performance characterizations. The hard-deadline downlink algorithm satisfies the average power constraint, the RT QoS constraints, and NRT queue stability, while achieving average NRT sum throughput within a bounded gap of optimal,
6
and simulations indicate average complexity close to linear in the number of RT users (Ewaisha et al., 2017). The cognitive-radio FH-AODPA results show fast SIR convergence, lower energy consumption than competing methods, Bellman residuals and terminal constraint errors close to zero, and robustness to channel variations without requiring explicit channel knowledge (Xu et al., 2013).
6. Conceptual distinctions and scope of the term
Several misconceptions are ruled out by the cited formulations. First, finite-horizon joint power allocation is not synonymous with greedy control. FHJPA explicitly values future energy availability through the Bellman recursion, whereas GA ignores future value and can therefore be suboptimal except in regimes where abundant source energy makes myopic control nearly sufficient (Tripathi et al., 4 Aug 2025).
Second, finite-horizon joint power allocation is not equivalent to a single metric or a single network topology. The objective may be expected throughput in a wirelessly powered uplink (Abad et al., 2019), NRT sum throughput under RT hard deadlines (Ewaisha et al., 2017), SIR tracking with power minimization under channel uncertainty (Xu et al., 2013), or average secrecy energy efficiency with joint source and jammer control (Tripathi et al., 4 Aug 2025). This suggests that FHJPA is best understood as a family of finite-horizon stochastic control formulations rather than a single universally fixed algorithm.
Third, “joint” does not require centralized global optimization over identical user roles. In one case, jointness refers to a stopping time and a power-allocation sequence (Abad et al., 2019); in another, to joint RT/NRT scheduling, power, and admission control (Ewaisha et al., 2017); in another, to coordinated PU/SU power control in coexistence mode (Xu et al., 2013); and in the secrecy setting, to the simultaneous choice of source transmit power and destination jamming power (Tripathi et al., 4 Aug 2025).
Finally, finite-horizon control does not preclude analytically tractable structure. The literature contains threshold rules, closed-form per-slot power expressions, Lambert-7 solutions, water-filling-like policies, and distributed adaptive dynamic programming. A plausible implication is that the main difficulty in FHJPA is not finite horizon per se, but the interaction between finite horizon and the particular physical constraints—battery causality, hard deadlines, interference coupling, or uncertain dynamics—that determine whether the solution is closed form, lookup-table based, or adaptively learned.