Papers
Topics
Authors
Recent
Search
2000 character limit reached

Infinite-Horizon Joint Power Allocation (IHJPA)

Updated 7 July 2026
  • IHJPA is a reinforcement learning model for joint source transmit and destination jamming power allocation in energy-harvesting, finite-battery wireless secrecy systems.
  • It employs a stationary policy via policy iteration to optimize long-term secrecy energy efficiency under stochastic channel, energy harvesting, and battery dynamics.
  • The formulation benchmarks against finite-horizon and greedy algorithms, highlighting trade-offs in computational complexity and secrecy metrics.

Searching arXiv for the cited IHJPA paper and closely related work to ground the article in current literature. arXiv search query: (Tripathi et al., 4 Aug 2025) OR "Secure Energy Efficient Wireless Transmission: A Finite v/s Infinite-Horizon RL Solution" OR "Joint Transmit and Jamming Power Optimization for Secrecy in Energy Harvesting Networks: A Reinforcement Learning Approach" OR "Infinite Horizon Optimal Transmission Power Control for Remote State Estimation over Fading Channels" Infinite-Horizon Joint Power Allocation (IHJPA) denotes an infinite-horizon reinforcement-learning formulation for jointly selecting source transmit power and destination jamming power in an energy-harvesting wireless secrecy system with finite batteries and full-duplex jamming, as introduced in "Secure Energy Efficient Wireless Transmission: A Finite v/s Infinite-Horizon RL Solution" (Tripathi et al., 4 Aug 2025). In that setting, the horizon is treated as infinite or unknown, the network lifetime is represented through a discount factor Γ\Gamma interpreted as the probability that the network remains operational in a slot, and the control objective is long-run secrecy energy efficiency (SEE) under Markovian channel, harvesting, and battery dynamics. Related infinite-horizon joint allocation formulations appear in secrecy-maximization models with random lifetime (Tripathi et al., 2024), in average-cost belief-state control for remote estimation over fading channels (Ren et al., 2016), and, more indirectly, in infinite-battery structural power-allocation results based on reverse multi-stage water-filling (Gong et al., 2013).

1. System model and controlled resources

IHJPA in (Tripathi et al., 4 Aug 2025) is defined on a three-node wireless system in which source SS sends data to destination DD, a passive eavesdropper EE attempts interception, and DD uses full-duplex capability to receive from SS while simultaneously transmitting a jamming signal toward EE. Both SS and DD harvest energy and store it in finite batteries. Time is slotted, k{0,1,,K1}k \in \{0,1,\dots,K-1\}, and the controller chooses in each slot the source transmit power SS0 and the destination jamming power SS1.

The power pair is constrained by current battery levels:

SS2

Battery evolution is energy-causality based and includes saturation at battery maximum. Energy harvesting at each node is modeled as an independent Bernoulli process: SS3 with SS4, and SS5 with SS6. This makes future resource availability stochastic and couples present actions to future feasibility.

The secrecy mechanism is produced by full-duplex jamming. Residual self-interference at SS7 is controlled by the SIC parameter SS8, where SS9 denotes no SIC and DD0 denotes perfect SIC. The resulting SINRs are

DD1

Accordingly, increasing DD2 improves secrecy by jamming DD3, but it can also degrade reception at DD4 through residual self-interference. The joint optimization of DD5 and DD6 is therefore intrinsic rather than incidental.

2. SEE objective and infinite-horizon MDP formulation

The per-slot secrecy rate in (Tripathi et al., 4 Aug 2025) is

DD7

with

DD8

The paper’s central performance metric is the average secrecy energy efficiency over DD9 slots,

EE0

Problem EE1 maximizes EE2 subject to the battery update laws and per-slot power feasibility. Because power decisions affect later battery states and later admissible actions, the problem is a sequential stochastic control problem.

For IHJPA, the system is formulated as an MDP whose state in slot EE3 is

EE4

The state therefore contains four quantized channel gains and the two battery levels. The action is the power pair

EE5

chosen from the feasible set

EE6

If EE7 contains EE8 discrete power levels, then the discretized action space has size EE9.

The transition law factors into channel evolution, energy harvesting, and battery evolution. The battery transition terms are deterministic indicators: they are DD0 if the battery update equations are satisfied and DD1 otherwise. The immediate reward is the per-slot SEE,

DD2

The infinite-horizon objective is tied to the discount factor DD3, interpreted as the probability that the system survives another slot. The paper characterizes this as the standard infinite-horizon RL setup in which the policy is stationary and the Bellman optimality principle is used (Tripathi et al., 4 Aug 2025).

3. Policy-iteration structure and stationary control law

IHJPA is an infinite-horizon policy-iteration-based solution. The method uses the same state and action spaces as the finite-horizon formulation, assumes an infinite or unknown horizon, computes a stationary optimal policy in a lookup table, and then follows that policy slot by slot during transmission. The paper states that it implements the planning phase of a prior RL method from the authors’ earlier work (Tripathi et al., 4 Aug 2025).

The closely related secrecy formulation in "Joint Transmit and Jamming Power Optimization for Secrecy in Energy Harvesting Networks: A Reinforcement Learning Approach" (Tripathi et al., 2024) makes the policy-iteration structure explicit. There, random lifetime DD4 with per-slot survival probability DD5 yields the discounted objective

DD6

and the Bellman equation is

DD7

Policy iteration alternates policy evaluation and policy improvement until policy stability; after planning, the policy is stored in a look-up table and used online. This is the same architectural pattern emphasized for IHJPA in (Tripathi et al., 4 Aug 2025): offline planning over a stationary MDP, followed by slot-wise execution through table lookup.

A common source of confusion is the meaning of the “infinite horizon.” In the secrecy MDP lineage represented by (Tripathi et al., 2024), the infinite-horizon model is not a literal claim that the network lives forever. Rather, the random stopping time is embedded as a discount factor DD8, and the value function sums future rewards with geometric discounting. In (Tripathi et al., 4 Aug 2025), the same interpretation appears through the probability that the network remains operational in a slot.

4. Relation to FHJPA and the greedy baseline

The paper positions IHJPA against two other decision rules: finite-horizon joint power allocation (FHJPA) and a greedy algorithm (GA). FHJPA is the finite-horizon optimal solution using backward induction. Starting from the final slot DD9, it computes the terminal action from immediate reward only and then recursively applies

SS0

The optimal action is stored as

SS1

By contrast, GA ignores future battery and channel effects and selects

SS2

Method Horizon/model Decision rule
FHJPA Known finite horizon Backward induction with stored optimal action
IHJPA Infinite or unknown horizon Stationary optimal policy via policy-iteration-style Bellman optimization
GA Myopic Maximizes immediate SEE

The substantive distinction is not merely algorithmic. FHJPA is optimal for a known finite horizon and explicitly accounts for the terminal structure of the transmission interval. IHJPA instead computes a stationary policy for an infinite or unknown horizon. GA is the least sophisticated formulation because it optimizes only the one-slot reward. The paper’s comparison shows that these distinctions matter operationally: FHJPA outperforms both GA and IHJPA in SEE because the transmission problem is modeled as finite horizon, while GA can approach FHJPA when the source node battery has sufficient energy (Tripathi et al., 4 Aug 2025).

5. Performance regime, metrics, and computational properties

IHJPA in (Tripathi et al., 4 Aug 2025) is evaluated against FHJPA and GA using three metrics: average SEE, expected total transmitted secure bits, and computational complexity. The expected secure-bit metric is

SS3

The reported ranking for SEE is that FHJPA performs best, IHJPA is worse when SS4 is small, and GA is below FHJPA but can approach it when source energy is abundant. The paper interprets this directly: FHJPA is superior because the problem is truly finite horizon, whereas IHJPA is a mismatch when the horizon is short.

The secure-bit comparison is more nuanced. FHJPA is not always best in secure bits because it intentionally trades throughput for SEE. Once SS5 is large enough, GA can outperform both FHJPA and IHJPA in secure bits. This establishes an important point of interpretation: maximizing SEE does not necessarily maximize total secure throughput. A common misconception is that an infinite-horizon joint policy optimized for energy efficiency must also dominate on aggregate secure bits; the reported results do not support that claim.

The effect of harvesting parameters is asymmetric. Increasing SS6 or SS7 generally improves SEE and secure bits until saturation. Increasing SS8 has a smaller effect on SEE, and source-side energy availability matters more than destination-side harvesting in this setup. The approximation regime of IHJPA is also explicit: as SS9 increases, the performance gap between FHJPA and IHJPA shrinks because the infinite-horizon model becomes more accurate for longer transmission horizons.

The paper reports the following complexity orders:

EE0

EE1

EE2

In the reported experiments, FHJPA takes EE3 percent less computational time than IHJPA (Tripathi et al., 4 Aug 2025). Thus, for the specific finite-horizon SEE problem studied there, IHJPA is neither the most accurate nor the least expensive option; its role is that of the infinite-horizon counterpart and approximation benchmark.

6. Broader research context and conceptual boundaries

The broader literature shows that IHJPA-type formulations are not confined to a single objective. In the secrecy setting of (Tripathi et al., 2024), the infinite-horizon MDP objective is long-term expected total secure bits transmitted until the network stops functioning, rather than SEE. That paper proposes an optimal joint power allocation (OJPA) algorithm based on policy iteration, alongside reduced-state and greedy approximations, and reports that the reduced-state variant lowers planning time by around EE4 in the tested setting. This indicates that the core infinite-horizon ingredients—Markov state evolution, stationary policies, and policy iteration—persist even when the reward changes from secrecy bits per energy to secrecy bits alone.

A second line of work, "Infinite Horizon Optimal Transmission Power Control for Remote State Estimation over Fading Channels" (Ren et al., 2016), places infinite-horizon joint power allocation in an average-cost belief-state MDP. There, the one-stage cost combines estimation error and power usage, and the central structural result is that an optimal deterministic and stationary policy exists. For scalar systems, the optimal transmission power is symmetric and monotonically increasing with respect to the innovation error. This is a different application domain, but it reinforces the general infinite-horizon principle that stationary control laws can be optimal when the problem is cast as an average-cost or discounted MDP.

A third related perspective appears in "Optimal Power Allocation for Energy Harvesting and Power Grid Coexisting Wireless Communication Systems" (Gong et al., 2013). That paper is formally a finite-horizon offline convex optimization problem, yet its infinite battery-capacity result induces a structure that behaves like an “infinite-horizon-like” offline policy. Reverse multi-stage water-filling yields non-decreasing water levels and a backward-recursive allocation rule. This is not an IHJPA formulation in the RL sense, but it is the closest analogue in that paper to an infinite-horizon structural policy.

Taken together, these results delimit the term with some precision. In its strict usage in (Tripathi et al., 4 Aug 2025), IHJPA is the stationary infinite-horizon RL solution for joint transmit and jamming power allocation under energy harvesting, finite batteries, and secrecy constraints. A plausible implication is that the term also points to a wider class of time-coupled joint allocation methods in which present power decisions are optimized through their effect on future state evolution rather than through immediate reward alone.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Infinite-Horizon Joint Power Allocation (IHJPA).