<?xml version="1.0" encoding="UTF-8"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/
         http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-10-08T21:51:02Z</responseDate><request verb="GetRecord" identifier="oai:lib.uoa.gr:uoadl:5358913" metadataPrefix="oai_dc">https://pergamos.lib.uoa.gr/uoa/dl/frontend/oaipmh</request><GetRecord><record>
      <header>
        <identifier>oai:lib.uoa.gr:uoadl:5358913</identifier>
        <datestamp>2026-03-05T11:51:42.180Z</datestamp>
        <setSpec>born_digital_postgraduate_thesis</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:dcterms="http://purl.org/dc/terms/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:identifier>uoadl:5358913</dc:identifier>
          <dc:identifier>https://pergamos.lib.uoa.gr/uoa/dl/object/5358913</dc:identifier>
          <dc:type>born_digital_postgraduate_thesis</dc:type>
          <dc:type xml:lang="el">Διπλωματική Εργασία</dc:type>
          <dc:type xml:lang="en">Postgraduate Thesis</dc:type>
          <dc:title>Contextual Bandits with Neural Networks and Trees</dc:title>
          <dc:date>2026</dc:date>
          <dc:language xml:lang="en">English</dc:language>
          <dc:creator xml:lang="el">ΚΩΝΣΤΑΝΤΑΤΟΣ ΚΩΝΣΤΑΝΤΙΝΟΣ</dc:creator>
          <dc:creator xml:lang="en">KONSTANTATOS KONSTANTINOS</dc:creator>
          <dc:description xml:lang="el">Πολλά προβλήματα περιλαμβάνουν τη λήψη διαδοχικών αποφάσεων υπό συν- θήκες αβεβαιότητας, όπου πρέπει να εξισορροπηθεί η εξερεύνηση (exploration) με την εκμετάλλευση (exploitation). Τα bandits παρέχουν ένα απλό μοντέλο για αυτό το δίλημμα. Τα contextual bandits αποτελούν μια πολύ σημαντική κατηγορία, όπου ο πράκτορας(agent) έχει πρόσβαση σε πρόσθετες πληροφορίες που μπορεί να βοηθήσουν στην πρόβλεψη της ποιότητας των ενεργειών του. Αυτό το πλαίσιο χρησιμοποιείται ευρέως σε πλήθος εφαρμογών. Η παρούσα διπλωματική εργασία παρέχει μια εισαγωγή στο πλαίσιο των bandits και επικεντρώνεται στα contextual bandits. Εισάγεται η βασική σημειογραφία που χρησιμοποιείται εκτενώς στη βιβλιογραφία, μαζί με αλγορίθμους, συμπερ- ιλαμβανομένων των ε-greedy, Upper Confidence Bound (UCB) και Thompson Sampling, οι οποίοι χρησιμοποιούνται για τη μείωση της μεταμέλειας (regret). Στη συνέχεια, η εργασία ανασκοπεί δύο από τις σημαντικότερες μεθό- δους στατιστικής μάθησης που χρησιμοποιούνται για τη μοντελοποίηση των σχέσεων μεταξύ πλαισίου και ανταμοιβής, συγκεκριμένα τα δέντρα απόφασης (και ensembles όπως τα τυχαία δάση) και τα νευρωνικά δίκτυα, συμπεριλαμ- βανομένου του Neural Tangent Kernel (NTK). Τέλος, η εργασία συνοψίζει τις βασικές ιδέες ορισμένων εμβληματικών ερ- γασιών, στις οποίες οι στατιστικές μέθοδοι αυτές χρησιμοποιήθηκαν για την προσέγγιση της συνάρτησης πλαισίου-ανταμοιβής στα contextual bandits.</dc:description>
          <dc:description xml:lang="en">Many problems involve sequential decision-making under uncertainty, where one must balance exploration and exploitation. Bandits provide a simple model for this dilemma. Contextual bandits is a very important category of bandits, where the agent has access to additional information that may help predict the quality of the actions. This framework is widely used in applications. This thesis provides an introduction to the bandit framework and focuses on contextual bandits. Basic notation that is used extensively in the bandit literature is introduced, along with standard algorithms including ε-greedy, Upper Confidence Bound, and Thompson Sampling which are used to reduce regret. The thesis then reviews two of the most important statistical learning methods used to model context–reward relationships, namely decision trees (and ensembles such as random forests) and neural networks, including the neural tangent kernel. The thesis also summarizes key ideas of some milestone papers, in which these methods were used to approximate the context-reward function in contextual bandits.</dc:description>
          <dc:subject xml:lang="el">Θετικές Επιστήμες</dc:subject>
          <dc:subject xml:lang="en">Science</dc:subject>
        </oai_dc:dc>
      </metadata>
    </record></GetRecord></OAI-PMH>