To see the other types of publications on this topic, follow the link: Offline Contextual Bandit.

Journal articles on the topic 'Offline Contextual Bandit'

Create a spot-on reference in APA, MLA, Chicago, Harvard, and other styles

Select a source type:

Consult the top 15 journal articles for your research on the topic 'Offline Contextual Bandit.'

Next to every source in the list of references, there is an 'Add to bibliography' button. Press on it, and we will generate automatically the bibliographic reference to the chosen work in the citation style you need: APA, MLA, Harvard, Chicago, Vancouver, etc.

You can also download the full text of the academic publication as pdf and read online its abstract whenever available in the metadata.

Browse journal articles on a wide variety of disciplines and organise your bibliography correctly.

1

Huang, Wen, and Xintao Wu. "Robustly Improving Bandit Algorithms with Confounded and Selection Biased Offline Data: A Causal Approach." Proceedings of the AAAI Conference on Artificial Intelligence 38, no. 18 (2024): 20438–46. http://dx.doi.org/10.1609/aaai.v38i18.30027.

Full text
Abstract:
This paper studies bandit problems where an agent has access to offline data that might be utilized to potentially improve the estimation of each arm’s reward distribution. A major obstacle in this setting is the existence of compound biases from the observational data. Ignoring these biases and blindly fitting a model with the biased data could even negatively affect the online learning phase. In this work, we formulate this problem from a causal perspective. First, we categorize the biases into confounding bias and selection bias based on the causal structure they imply. Next, we extract the
APA, Harvard, Vancouver, ISO, and other styles
2

Narita, Yusuke, Shota Yasui, and Kohei Yata. "Efficient Counterfactual Learning from Bandit Feedback." Proceedings of the AAAI Conference on Artificial Intelligence 33 (July 17, 2019): 4634–41. http://dx.doi.org/10.1609/aaai.v33i01.33014634.

Full text
Abstract:
What is the most statistically efficient way to do off-policy optimization with batch data from bandit feedback? For log data generated by contextual bandit algorithms, we consider offline estimators for the expected reward from a counterfactual policy. Our estimators are shown to have lowest variance in a wide class of estimators, achieving variance reduction relative to standard estimators. We then apply our estimators to improve advertisement design by a major advertisement company. Consistent with the theoretical result, our estimators allow us to improve on the existing bandit algorithm w
APA, Harvard, Vancouver, ISO, and other styles
3

Seifi, Farshad, and Seyed Taghi Akhavan Niaki. "Optimizing contextual bandit hyperparameters: A dynamic transfer learning-based framework." International Journal of Industrial Engineering Computations 15, no. 4 (2024): 951–64. http://dx.doi.org/10.5267/j.ijiec.2024.6.003.

Full text
Abstract:
The stochastic contextual bandit problem, recognized for its effectiveness in navigating the classic exploration-exploitation dilemma through ongoing player-environment interactions, has found broad applications across various industries. This utility largely stems from the algorithms’ ability to accurately forecast reward functions and maintain an optimal balance between exploration and exploitation, contingent upon the precise selection and calibration of hyperparameters. However, the inherently dynamic and real-time nature of bandit environments significantly complicates hyperparameter tuni
APA, Harvard, Vancouver, ISO, and other styles
4

Krishnamurthy, Sanath Kumar, Tanmay Gangwani, Sumeet Katariya, Branislav Kveton, Shrey Modi, and Anshuka Rangi. "Selective Uncertainty Propagation in Offline RL." Proceedings of the AAAI Conference on Artificial Intelligence 39, no. 17 (2025): 17974–82. https://doi.org/10.1609/aaai.v39i17.33977.

Full text
Abstract:
We consider the finite-horizon offline reinforcement learning (RL) setting, and are motivated by the challenge of learning the policy at any step h in dynamic programming (DP) algorithms. To learn this, it is sufficient to evaluate the treatment effect of deviating from the behavioral policy at step h after having optimized the policy for all future steps. Since the policy at any step can affect next-state distributions, the related distributional shift challenges can make this problem far more statistically hard than estimating such treatment effects in the stochastic contextual bandit settin
APA, Harvard, Vancouver, ISO, and other styles
5

Degroote, Hans, Patrick De Causmaecker, Bernd Bischl, and Lars Kotthoff. "A Regression-Based Methodology for Online Algorithm Selection." Proceedings of the International Symposium on Combinatorial Search 9, no. 1 (2021): 37–45. http://dx.doi.org/10.1609/socs.v9i1.18458.

Full text
Abstract:
Algorithm selection approaches have achieved impressive performance improvements in many areas of AI. Most of the literature considers the offline algorithm selection problem, where the initial selection model is never updated after training. However, new data from running algorithms on instances becomes available when algorithms are selected and run. We investigate how this online data can be used to improve the selection model over time. This is especially relevant when insufficient training instances were used, but potentially improves the performance of algorithm selection in all cases. We
APA, Harvard, Vancouver, ISO, and other styles
6

Li, Zhao, Junshuai Song, Zehong Hu, Zhen Wang, and Jun Gao. "Constrained Dual-Level Bandit for Personalized Impression Regulation in Online Ranking Systems." ACM Transactions on Knowledge Discovery from Data 16, no. 2 (2021): 1–23. http://dx.doi.org/10.1145/3461340.

Full text
Abstract:
Impression regulation plays an important role in various online ranking systems, e.g. , e-commerce ranking systems always need to achieve local commercial demands on some pre-labeled target items like fresh item cultivation and fraudulent item counteracting while maximizing its global revenue. However, local impression regulation may cause “butterfly effects” on the global scale, e.g. , in e-commerce, the price preference fluctuation in initial conditions (overpriced or underpriced items) may create a significantly different outcome, thus affecting shopping experience and bringing economic los
APA, Harvard, Vancouver, ISO, and other styles
7

Bhatt, Umang, Valerie Chen, Katherine M. Collins, et al. "Learning Personalized Decision Support Policies." Proceedings of the AAAI Conference on Artificial Intelligence 39, no. 13 (2025): 14203–11. https://doi.org/10.1609/aaai.v39i13.33555.

Full text
Abstract:
Individual human decision-makers may benefit from different forms of support to improve decision outcomes, but when will each form of support yield better outcomes? In this work, we posit that personalizing access to decision support tools can be an effective mechanism for instantiating the appropriate use of AI assistance. Specifically, we propose the general problem of learning a decision support policy that, for a given input, chooses which form of support to provide to decision-makers for whom we initially have no prior information. We develop Modiste, an interactive tool to learn personal
APA, Harvard, Vancouver, ISO, and other styles
8

Vera, Alberto, Siddhartha Banerjee, and Itai Gurvich. "Online Allocation and Pricing: Constant Regret via Bellman Inequalities." Operations Research 69, no. 3 (2021): 821–40. http://dx.doi.org/10.1287/opre.2020.2061.

Full text
Abstract:
We develop a framework for designing simple and efficient policies for a family of online allocation and pricing problems that includes online packing, budget-constrained probing, dynamic pricing, and online contextual bandits with knapsacks. In each case, we evaluate the performance of our policies in terms of their regret (i.e., additive gap) relative to an offline controller that is endowed with more information than the online controller. Our framework is based on Bellman inequalities, which decompose the loss of an algorithm into two distinct sources of error: (1) arising from computation
APA, Harvard, Vancouver, ISO, and other styles
9

Aditya Kambhampati. "Advances in Personalized Investment Advisory through Reinforcement Learning: A Technical Review." Journal of Computer Science and Technology Studies 7, no. 4 (2025): 187–93. https://doi.org/10.32996/jcsts.2025.7.4.22.

Full text
Abstract:
Reinforcement learning (RL) represents a transformative technology in personalized investment advisory services, addressing fundamental limitations of traditional static approaches. This article explores the application of diverse RL frameworks to financial decision-making, from contextual multi-armed bandits for tactical allocations to full Markov Decision Processes for long-term planning. The integration of sophisticated state representations, multi-objective reward functions, and offline learning methodologies enables systems that adapt to individual investor behaviors while maintaining app
APA, Harvard, Vancouver, ISO, and other styles
10

Ayle, Morgane, Jimmy Tekli, Julia El-Zini, Boulos El-Asmar, and Mariette Awad. "BAR — A Reinforcement Learning Agent for Bounding-Box Automated Refinement." Proceedings of the AAAI Conference on Artificial Intelligence 34, no. 03 (2020): 2561–68. http://dx.doi.org/10.1609/aaai.v34i03.5639.

Full text
Abstract:
Research has shown that deep neural networks are able to help and assist human workers throughout the industrial sector via different computer vision applications. However, such data-driven learning approaches require a very large number of labeled training images in order to generalize well and achieve high accuracies that meet industry standards. Gathering and labeling large amounts of images is both expensive and time consuming, specifically for industrial use-cases. In this work, we introduce BAR (Bounding-box Automated Refinement), a reinforcement learning agent that learns to correct ina
APA, Harvard, Vancouver, ISO, and other styles
11

Simchi-Levi, David, and Yunzong Xu. "Bypassing the Monster: A Faster and Simpler Optimal Algorithm for Contextual Bandits Under Realizability." Mathematics of Operations Research, December 9, 2021. http://dx.doi.org/10.1287/moor.2021.1193.

Full text
Abstract:
We consider the general (stochastic) contextual bandit problem under the realizability assumption, that is, the expected reward, as a function of contexts and actions, belongs to a general function class [Formula: see text]. We design a fast and simple algorithm that achieves the statistically optimal regret with only [Formula: see text] calls to an offline regression oracle across all T rounds. The number of oracle calls can be further reduced to [Formula: see text] if T is known in advance. Our results provide the first universal and optimal reduction from contextual bandits to offline regre
APA, Harvard, Vancouver, ISO, and other styles
12

Oztop, Erhan, Suzan Ece Ada, and Emre Ugur. "Diffusion Policies for Out-of-Distribution Generalization in Offline Reinforcement Learning." February 7, 2024. https://doi.org/10.48550/arXiv.2307.04726.

Full text
Abstract:
Offline Reinforcement Learning (RL) methods leverage previous experiences to learn better policies than the behavior policy used for data collection. In contrast to behavior cloning, which assumes the data is collected from expert demonstrations, offline RL can work with non-expert data and multimodal behavior policies. However, offline RL algorithms face challenges in handling distribution shifts and effectively representing policies due to the lack of online interaction during training. Prior work on offline RL uses conditional diffusion models to represent multimodal behavior in the dataset
APA, Harvard, Vancouver, ISO, and other styles
13

Soemers, Dennis, Tim Brys, Kurt Driessens, Mark Winands, and Ann Nowé. "Adapting to Concept Drift in Credit Card Transaction Data Streams Using Contextual Bandits and Decision Trees." Proceedings of the AAAI Conference on Artificial Intelligence 32, no. 1 (2018). http://dx.doi.org/10.1609/aaai.v32i1.11411.

Full text
Abstract:
Credit card transactions predicted to be fraudulent by automated detection systems are typically handed over to human experts for verification. To limit costs, it is standard practice to select only the most suspicious transactions for investigation. We claim that a trade-off between exploration and exploitation is imperative to enable adaptation to changes in behavior (concept drift). Exploration consists of the selection and investigation of transactions with the purpose of improving predictive models, and exploitation consists of investigating transactions detected to be suspicious. Modelin
APA, Harvard, Vancouver, ISO, and other styles
14

Cao, Junyu, and Wei Sun. "Tiered Assortment: Optimization and Online Learning." Management Science, October 4, 2023. http://dx.doi.org/10.1287/mnsc.2023.4940.

Full text
Abstract:
Due to the sheer number of available choices, online retailers frequently use tiered assortment to present products. In this case, groups of products are arranged across multiple pages or stages, and a customer clicks on “next” or “load more” to access them sequentially. Despite the prevalence of such assortments in practice, this topic has not received much attention in the existing literature. In this work, we focus on a sequential choice model that captures customers’ behavior when product recommendations are presented in tiers. We analyze multiple variants of tiered assortment optimization
APA, Harvard, Vancouver, ISO, and other styles
15

Zeng, Yingyan, Xiaoyu Chen, and Ran Jin. "Ensemble Active Learning by Contextual Bandits for AI Incubation in Manufacturing." ACM Transactions on Intelligent Systems and Technology, October 25, 2023. http://dx.doi.org/10.1145/3627821.

Full text
Abstract:
An Industrial Cyber-physical System (ICPS) provide a digital foundation for data-driven decision-making by artificial intelligence (AI) models. However, the poor data quality (e.g., inconsistent distribution, imbalanced classes) of high-speed, large-volume data streams poses significant challenges to the online deployment of offline-trained AI models. As an alternative, updating AI models online based on streaming data enables continuous improvement and resilient modeling performance. However, for a supervised learning model (i.e., a base learner), it is labor-intensive to annotate all streami
APA, Harvard, Vancouver, ISO, and other styles
We offer discounts on all premium plans for authors whose works are included in thematic literature selections. Contact us to get a unique promo code!