{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaL4","dataSources":[{"sourceId":84795,"databundleVersionId":10462807,"sourceType":"competition"}],"dockerImageVersionId":30805,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<head>\n  <!-- Add Google Font link -->\n  <link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@400;700&display=swap\" rel=\"stylesheet\">\n</head>\n\n<div style=\"border: 5px solid #f39c12; border-radius: 10px; padding: 15px; text-align: left; font-family: 'Roboto', sans-serif; width: 80%; max-width: 700px; margin: auto; background-color: #2c3e50; color: white;\">\n  <h1 style=\"background-color: #e74c3c; padding: 10px; border-radius: 5px; text-align: center; font-size: 1.8em;\">Titanic - Machine Learning from Disaster </h1>\n  \n  <h4>Introduction</h4>\n  <ul>\n    <li>The Konwinski Prize competition focuses on applying automated inference techniques to software development issues, specifically GitHub repositories. The task is to develop a system that can analyze a problem statement related to a bug or enhancement and then automatically generate a patch (in the form of a diff) to fix the codebase in a GitHub repository. This involves understanding both the code in the repository and the problem described in the issue, and then creating an appropriate fix.</li>\n  </ul>\n\n  <h4>Goal</h4>\n  <ul>\n    <li>\r\n\r\nThe goal of the Konwinski Prize competition is to build an automated system that can analyze a bug or enhancement description (problem statement) and a corresponding GitHub repository, and then generate an appropriate code patch to fix the issue described.</li>\n  </ul>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# **Step 1: Set Up the Environment**","metadata":{}},{"cell_type":"code","source":"\nimport io\nimport os\nimport shutil\n\nimport pandas as pd\nimport polars as pl\n\nimport kaggle_evaluation.konwinski_prize_inference_server","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T17:01:54.310882Z","iopub.execute_input":"2024-12-12T17:01:54.31166Z","iopub.status.idle":"2024-12-12T17:01:54.31511Z","shell.execute_reply.started":"2024-12-12T17:01:54.311626Z","shell.execute_reply":"2024-12-12T17:01:54.314447Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **Step 2: Define the Inference Server**","metadata":{}},{"cell_type":"code","source":"\"\"\"\n# Set up the Inference Server\ninstance_count = None\n\ndef get_number_of_instances(num_instances: int) -> None:\n  \n    global instance_count\n    instance_count = num_instances\n\nfirst_prediction = True\n\ndef predict(problem_statement: str, repo_archive: io.BytesIO) -> str:\n \n    global first_prediction\n    if not first_prediction:\n        return None  # Skip after the first prediction.\n\n    # Unpack the repository\n    with open('repo_archive.tar', 'wb') as f:\n        f.write(repo_archive.read())\n    repo_path = 'repo'\n    if os.path.exists(repo_path):\n        shutil.rmtree(repo_path)\n    shutil.unpack_archive('repo_archive.tar', extract_dir=repo_path)\n    os.remove('repo_archive.tar')\n    first_prediction = False\n    \n    # Load pre-trained model for patch generation\n    model = T5ForConditionalGeneration.from_pretrained('t5-base')\n    tokenizer = T5Tokenizer.from_pretrained('t5-base')\n\n    # Format the input for the model\n    input_text = f\"Fix this issue: {problem_statement}\"\n    inputs = tokenizer(input_text, return_tensors=\"pt\", max_length=512, truncation=True, padding=True)\n\n    # Generate the patch\n    with torch.no_grad():\n        output = model.generate(inputs['input_ids'], max_length=512, num_beams=4, early_stopping=True)\n\n    # Decode the generated patch\n    patch = tokenizer.decode(output[0], skip_special_tokens=True)\n\n    return patch\n    \"\"\"\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T17:01:54.338022Z","iopub.execute_input":"2024-12-12T17:01:54.33864Z","iopub.status.idle":"2024-12-12T17:01:54.343332Z","shell.execute_reply.started":"2024-12-12T17:01:54.33861Z","shell.execute_reply":"2024-12-12T17:01:54.342712Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Global variables\ninstance_count = None\n\n# This function receives the total number of instances to be served.\ndef get_number_of_instances(num_instances: int) -> None:\n    \"\"\"The very first message from the gateway will be the total number of instances to be served.\"\"\"\n    global instance_count\n    instance_count = num_instances\n\n\n# This will be the first prediction flag\nfirst_prediction = True\n\ndef predict(problem_statement: str, repo_archive: io.BytesIO) -> str:\n    \"\"\" Replace this function with your inference code.\n    Args:\n        problem_statement: The text of the git issue.\n        repo_archive: A BytesIO buffer path with a .tar containing the codebase that must be patched.\n    \"\"\"\n    global first_prediction\n    \n    if not first_prediction:\n        return None  # Skip issue. We can apply logic to skip inference if needed.\n\n    # Write the repo archive to a file\n    with open('repo_archive.tar', 'wb') as f:\n        f.write(repo_archive.read())\n    \n    # Set up the path to extract the repository\n    repo_path = 'repo'\n    \n    # If the repo path exists, remove it\n    if os.path.exists(repo_path):\n        shutil.rmtree(repo_path)\n    \n    # Unpack the .tar archive into the repo directory\n    shutil.unpack_archive('repo_archive.tar', extract_dir=repo_path)\n    \n    # Remove the .tar file after extracting\n    os.remove('repo_archive.tar')\n    \n    first_prediction = False  # Set this flag to False to skip future predictions\n\n    # Instead of doing actual diffing or patching logic, we return a simple string for now\n    # This will definitely fail, but you can replace it with your actual inference logic.\n    return \"Hello World\"\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T17:01:54.356882Z","iopub.execute_input":"2024-12-12T17:01:54.357467Z","iopub.status.idle":"2024-12-12T17:01:54.362711Z","shell.execute_reply.started":"2024-12-12T17:01:54.357432Z","shell.execute_reply":"2024-12-12T17:01:54.362106Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **Step 3: Load and Explore the Dataset**","metadata":{}},{"cell_type":"code","source":"import zipfile\nimport os\n\n# Path to the input directory\ninput_dir = '/kaggle/input/konwinski-prize/'  # Path to the input directory\n\n# Path to the zipped dataset\nzip_file_path = os.path.join(input_dir, 'data.a_zip')  # Path to the .zip file containing data\n\n# Define the extraction directory\nextracted_dir = '/kaggle/tmp/extracted_data/'  # Directory to store extracted data\n\n# Ensure the extraction directory exists\nos.makedirs(extracted_dir, exist_ok=True)\n\n# Extract the zip file\nwith zipfile.ZipFile(zip_file_path, 'r') as zip_ref:\n    zip_ref.extractall(extracted_dir)\n\nprint(f\"Data has been extracted to: {extracted_dir}\")\n\n# List the files to inspect what was extracted\nextracted_files = os.listdir(extracted_dir)\nextracted_files\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T17:01:54.371838Z","iopub.execute_input":"2024-12-12T17:01:54.372419Z","iopub.status.idle":"2024-12-12T17:01:55.867216Z","shell.execute_reply.started":"2024-12-12T17:01:54.372392Z","shell.execute_reply":"2024-12-12T17:01:55.86657Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import polars as pl\nimport pandas as pd\n\n# Load the train metadata (data.parquet) for inspection\ndata_parquet_path = os.path.join(extracted_dir, 'data/data.parquet')\n\n# Use Polars for efficient handling of large datasets\ndata_t= pl.read_parquet(data_parquet_path)\n# Convert Polars DataFrame to Pandas DataFrame\ndata = data_t.to_pandas()\n\n# Show the first few rows of the dataset\ndata.head()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T17:01:55.868408Z","iopub.execute_input":"2024-12-12T17:01:55.868681Z","iopub.status.idle":"2024-12-12T17:01:55.891472Z","shell.execute_reply.started":"2024-12-12T17:01:55.868655Z","shell.execute_reply":"2024-12-12T17:01:55.890832Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **Step 4: Preprocess the Data**","metadata":{}},{"cell_type":"code","source":"\"\"\"# Prepare data for model training (tokenize the problem statement and patch)\ninput_texts = train_data['problem_statement'].values\noutput_texts = train_data['patch'].values\n\n# Create a tokenizer using a pre-trained T5 model\ntokenizer = T5Tokenizer.from_pretrained('t5-base')\n\ndef tokenize_data(input_texts, output_texts):\n    # Tokenize the data\n    inputs = tokenizer(input_texts.tolist(), max_length=512, padding=True, truncation=True, return_tensors=\"pt\")\n    outputs = tokenizer(output_texts.tolist(), max_length=512, padding=True, truncation=True, return_tensors=\"pt\")\n\n    return inputs, outputs\n\ninputs, outputs = tokenize_data(input_texts, output_texts)\"\"\"\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T17:01:55.892326Z","iopub.execute_input":"2024-12-12T17:01:55.89282Z","iopub.status.idle":"2024-12-12T17:01:56.468475Z","shell.execute_reply.started":"2024-12-12T17:01:55.892791Z","shell.execute_reply":"2024-12-12T17:01:56.467788Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **Step 5: Train the Model**","metadata":{}},{"cell_type":"code","source":"\"\"\"# Load the T5 model for sequence-to-sequence learning\nmodel = T5ForConditionalGeneration.from_pretrained('t5-base')\n\n# Create a TensorDataset for training\nfrom torch.utils.data import DataLoader, TensorDataset\n\n# Create a DataLoader to handle batching\ntrain_dataset = TensorDataset(inputs['input_ids'], outputs['input_ids'], outputs['attention_mask'])\ntrain_dataloader = DataLoader(train_dataset, batch_size=8, shuffle=True)\n\n# Optimizer setup\nfrom torch.optim import AdamW\n\noptimizer = AdamW(model.parameters(), lr=5e-5)\nmodel.to('cuda')  # Move model to GPU if available\n\n# Training loop (simplified)\nepochs = 3\nfor epoch in range(epochs):\n    model.train()\n    for batch in train_dataloader:\n        input_ids, target_ids, attention_mask = batch\n        input_ids = input_ids.to('cuda')\n        target_ids = target_ids.to('cuda')\n        attention_mask = attention_mask.to('cuda')\n\n        optimizer.zero_grad()\n        output = model(input_ids=input_ids, labels=target_ids, attention_mask=attention_mask)\n        loss = output.loss\n        loss.backward()\n        optimizer.step()\n\n    print(f\"Epoch {epoch+1}/{epochs}, Loss: {loss.item()}\")\n\"\"\"","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T17:01:56.469899Z","iopub.execute_input":"2024-12-12T17:01:56.470401Z","iopub.status.idle":"2024-12-12T17:01:56.477626Z","shell.execute_reply.started":"2024-12-12T17:01:56.470371Z","shell.execute_reply":"2024-12-12T17:01:56.477025Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **Step 6: Save the Trained Model**","metadata":{}},{"cell_type":"code","source":"\"\"\"# Save the trained model\nmodel.save_pretrained('./patched_model')\ntokenizer.save_pretrained('./patched_model')\n\"\"\"","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T17:01:56.478435Z","iopub.execute_input":"2024-12-12T17:01:56.478691Z","iopub.status.idle":"2024-12-12T17:01:56.488425Z","shell.execute_reply.started":"2024-12-12T17:01:56.478667Z","shell.execute_reply":"2024-12-12T17:01:56.487853Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **Step 7: Run the Inference Server Locally**","metadata":{}},{"cell_type":"code","source":"\n# Set up the inference server\ninference_server = kaggle_evaluation.konwinski_prize_inference_server.KPrizeInferenceServer(\n    get_number_of_instances,   \n    predict\n)\n\n# Run the inference server locally (during competition, this will run on Kaggle)\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n    inference_server.serve()\nelse:\n    inference_server.run_local_gateway(\n        data_paths=(\n            '/kaggle/input/konwinski-prize/',  # Path to the competition data\n            '/kaggle/tmp/konwinski-prize/',   # Path for temporary unpacking\n        )\n    )\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-12T17:01:56.489189Z","iopub.execute_input":"2024-12-12T17:01:56.489634Z","iopub.status.idle":"2024-12-12T17:02:14.487032Z","shell.execute_reply.started":"2024-12-12T17:01:56.489608Z","shell.execute_reply":"2024-12-12T17:02:14.48612Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"border: 2px solid #FFA500; border-radius: 10px; padding: 10px; background-color: #FFF5E6; text-align: center; font-family: Arial, sans-serif; width: 80%; max-width: 600px; margin: auto;\">\n  <h3 style=\"color: #FFA500;\">👍 <strong>Enjoyed this guide?</strong></h3>\n  <p style=\"color: #333333;\">If you found this guide helpful, please consider giving it an upvote! Your support helps us continue to create valuable content and improve our resources.</p>\n  <p style=\"font-size: 16px; color: #FF8C00;\">Thank you! 😊</p>\n</div>","metadata":{}},{"cell_type":"markdown","source":"\n### You can connect with me on : [Linkedin](https://www.linkedin.com/in/m-abdullah-2a1631315/)","metadata":{}}]}