{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84795,"databundleVersionId":10462807,"sourceType":"competition"},{"sourceId":208761,"sourceType":"modelInstanceVersion","isSourceIdPinned":true,"modelInstanceId":177918,"modelId":200222}],"dockerImageVersionId":30822,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 1. Initial Steps & Explore the Model page here \n   - Include this \"andro_konwinski_llm_model_readiness\" model\n   - moving .py into kaggle/working folder since /kaggle/input/ folder is read only path\n   - check here usage of this ml model dataset readiness [https://www.kaggle.com/models/andrometocs/andro_konwinski_llm_model_readiness/competitions](https://www.kaggle.com/models/andrometocs/andro_konwinski_llm_model_readiness/competitions)\n   - Follow more here for Pretrained ML Models, Utility for ML Model datasets readiness and datasets\n     [Andrometocs - ML Models and Datasets ](https://www.kaggle.com/organizations/andrometocs)","metadata":{}},{"cell_type":"markdown","source":"\n### Code Snippet 1: Copying Files\n```bash\n!cp -r /kaggle/input/andro_konwinski_llm_model_readiness/other/default/2/* /kaggle/working\n```\nThis command is a shell command that runs in a Jupyter notebook (or a Kaggle notebook in this case). Here’s what each part of this command does:\n\n- `!`: The exclamation mark is used to run shell commands directly from a Jupyter notebook cell.\n- `cp`: The `cp` command is used to copy files and directories.\n- `-r`: The `-r` flag stands for \"recursive,\" meaning it will copy all files and directories within the specified directory.\n- `/kaggle/input/andro_konwinski_llm_model_readiness/other/default/1/*`: This is the source directory. The asterisk `*` means \"all files and directories within this directory.\"\n- `/kaggle/working`: This is the destination directory where the files will be copied to.\n\n**In summary:** This command copies all files and directories from `/kaggle/input/andro_konwinski_llm_model_readiness/other/default/1/` to the `/kaggle/working` directory.\n\n### Code Snippet 2: Modifying the Python Path\n```python\nimport sys \nsys.path.append('/kaggle/working/andro_swebench_processor.py')\n```\nThis Python code snippet modifies the Python path to include the specified directory. Here’s what each part does:\n\n- `import sys`: This imports the `sys` module, which provides access to some variables used or maintained by the interpreter and functions that interact with the interpreter.\n- `sys.path.append('/kaggle/working/andro_swebench_processor.py')`: This line adds the specified directory (`'/kaggle/working/andro_swebench_processor.py'`) to the Python path. By appending this path, Python will be able to import modules from this directory.\n\n**In summary:** This code snippet ensures that the Python interpreter can find and import the `andro_swebench_processor.py` script, which is now located in the `/kaggle/working` directory.\n\n### Combined Explanation\nTogether, these snippets:\n1. **Copy** all files from a specified source directory to the working directory in the Kaggle environment.\n2. **Modify the Python path** to include the copied script, making it available for import and use in the notebook.","metadata":{}},{"cell_type":"code","source":"!cp -r /kaggle/input/andro_konwinski_llm_model_readiness/other/default/2/* /kaggle/working\nimport sys \nsys.path.append('/kaggle/working/andro_swebench_processor.py')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-24T13:22:53.012342Z","iopub.execute_input":"2024-12-24T13:22:53.0126Z","iopub.status.idle":"2024-12-24T13:22:53.138605Z","shell.execute_reply.started":"2024-12-24T13:22:53.012572Z","shell.execute_reply":"2024-12-24T13:22:53.137451Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 2. Now EDA and Visualization, LLM QNA Dataset created","metadata":{"execution":{"iopub.status.busy":"2024-12-24T11:00:36.86568Z","iopub.execute_input":"2024-12-24T11:00:36.86613Z","iopub.status.idle":"2024-12-24T11:00:36.872532Z","shell.execute_reply.started":"2024-12-24T11:00:36.866098Z","shell.execute_reply":"2024-12-24T11:00:36.871227Z"}}},{"cell_type":"markdown","source":"\n### Step-by-Step Explanation\n\n1. **Import the `andro_swebench_processor` class**:\n   ```python\n   from andro_swebench_processor import andro_swebench_processor\n   ```\n   This imports the `andro_swebench_processor` class from the `andro_swebench_processor` module, making it available for use in the notebook.\n\n2. **Initialize the `processor` object**:\n   ```python\n   processor = andro_swebench_processor(\n       zip_file_path='../input/konwinski-prize/data.a_zip',\n       parquet_file_path='data/data.parquet',\n       dataset_name='princeton-nlp/SWE-bench'\n   )\n   ```\n   This initializes an instance of the `andro_swebench_processor` class with the specified arguments:\n   - `zip_file_path`: Path to the zip file containing data.\n   - `parquet_file_path`: Path to the parquet file inside the zip.\n   - `dataset_name`: Name of the SWE-bench dataset.\n\n3. **Visualize the data**:\n   ```python\n   print(\"Visualizing data...\")\n   processor.visualize_data()\n   ```\n   This calls the `visualize_data` method of the `processor` object to generate visualizations (e.g., word clouds, statement length distributions).\n\n4. **Create the QNA dataset as a list**:\n   ```python\n   print(\"Creating QNA dataset as list...\")\n   processor.create_QNA_dataset_as_list()\n   ```\n   This calls the `create_QNA_dataset_as_list` method to create a QNA dataset in the form of a list of formatted question-answer templates.\n\n5. **Store the QNA dataset list**:\n   ```python\n   qna_dataset_list = processor.get_all_QNA_dataset()\n   ```\n   This stores the generated QNA dataset list in the `qna_dataset_list` variable.\n\n6. **Show the size of the QNA list**:\n   ```python\n   print(\"QNA List Size 1121:\")\n   processor.show_list_size()\n   ```\n   This calls the `show_list_size` method to display the size of the QNA dataset list.\n\n7. **Create the QNA dataset as a DataFrame**:\n   ```python\n   print(\"Creating QNA dataset as DataFrame...\")\n   processor.create_QNA_dataset_as_dataframe()\n   ```\n   This calls the `create_QNA_dataset_as_dataframe` method to create a QNA dataset in the form of a pandas DataFrame.\n\n8. **Get and display the DataFrame summary**:\n   ```python\n   print(\"DataFrame Summary:\")\n   processor.get_dataframe_summary()\n   ```\n   This calls the `get_dataframe_summary` method to display the shape, info, and description of the QNA DataFrame.\n\n9. **Print a sample of the QNA dataset**:\n   ```python\n   print(\"Sample QNA dataset:\")\n   processor.print_colored_QNA_samples()\n   ```\n   This calls the `print_colored_QNA_samples` method to print samples from the QNA dataset with the question and answer in different colors.\n\n10. **Print the first QNA template from the list**:\n    ```python\n    print(\"First QNA template from the list:\")\n    print(qna_dataset_list[0] if qna_dataset_list else \"No data available.\")\n    ```\n    This prints the first QNA template from the list, if available, or a message indicating no data is available.\n\n### Summary\n\nThe provided code initializes an `andro_swebench_processor` object, processes a dataset from a zip file, generates visualizations, creates a QNA dataset in both list and DataFrame formats, displays summaries, and prints sample entries from the dataset. This setup is useful for processing and visualizing data in the context of a question-and-answer dataset.","metadata":{}},{"cell_type":"code","source":"    from andro_swebench_processor import andro_swebench_processor\n\n    processor = andro_swebench_processor(\n                        zip_file_path='../input/konwinski-prize/data.a_zip',\n                        parquet_file_path='data/data.parquet',\n                        dataset_name='princeton-nlp/SWE-bench'\n                    )\n\n    print(\"Visualizing data...\")\n    processor.visualize_data()\n\n    print(\"Creating QNA dataset as list...\")\n    processor.create_QNA_dataset_as_list()\n\n    # Store the QNA dataset list\n    qna_dataset_list = processor.get_all_QNA_dataset()\n\n    print(\"QNA List Size:\")\n    processor.show_list_size()\n\n    print(\"Creating QNA dataset as DataFrame...\")\n    processor.create_QNA_dataset_as_dataframe()\n\n    print(\"DataFrame Summary:\")\n    processor.get_dataframe_summary()\n\n    print(\"Sample QNA dataset:\")\n    processor.print_colored_QNA_samples()\n\n    # Example: Print the first QNA template from the list\n    print(\"First QNA template from the list:\")\n    print(qna_dataset_list[0] if qna_dataset_list else \"No data available.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T13:22:53.139592Z","iopub.execute_input":"2024-12-24T13:22:53.139853Z","iopub.status.idle":"2024-12-24T13:23:19.94391Z","shell.execute_reply.started":"2024-12-24T13:22:53.13983Z","shell.execute_reply":"2024-12-24T13:23:19.942871Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 3. Now QNA dataset is ready for LLM data model !!","metadata":{}},{"cell_type":"markdown","source":"# 4. Possible Choices to explore in ML, NLP, LLM, SLM","metadata":{}},{"cell_type":"markdown","source":"## 4.1. QNA Dataset ML/LLM/NLP \n\nCreating a QNA dataset list involves several stages of machine learning (ML), natural language processing (NLP), and large language models (LLMs), tailored for training a model capable of summarizing GitHub repositories. Let's break down the components and methods involved in this process:\n\n### Machine Learning (ML) Techniques\n1. **Supervised Learning**: The training process is typically supervised, using labeled data (QNA pairs) to teach the model how to generate summaries.\n2. **Transfer Learning**: Utilizing pre-trained models and fine-tuning them on specific tasks, such as summarization.\n\n### Natural Language Processing (NLP) Techniques\n1. **Tokenization**: Breaking down text into tokens (words, subwords, or characters) that the model can process.\n2. **Text Preprocessing**: Cleaning the text data by removing special characters, stop words, and normalizing the text.\n3. **Named Entity Recognition (NER)**: Identifying and classifying entities (such as repository names, function names) in the text.\n4. **Part-of-Speech Tagging (POS)**: Analyzing the grammatical structure of the text to understand its context.\n5. **Word Embeddings**: Converting words into numerical vectors that capture their semantic meaning, often using methods like Word2Vec, GloVe, or BERT embeddings.\n\n### Large Language Models (LLMs) and Specialized Language Models (SLMs)\n1. **BERT (Bidirectional Encoder Representations from Transformers)**: A transformer-based model pre-trained on a large corpus of text, fine-tuned for tasks like summarization and question-answering.\n2. **GPT (Generative Pre-trained Transformer)**: A transformer-based generative model capable of producing coherent text, often used for generating summaries and QNA pairs.\n3. **T5 (Text-To-Text Transfer Transformer)**: A transformer model that converts all NLP tasks into a text-to-text format, suitable for tasks like summarization and question-answer generation.\n4. **BART (Bidirectional and Auto-Regressive Transformers)**: A denoising autoencoder for pretraining sequence-to-sequence models, effective for text generation tasks including summarization.\n\n### Specialized Techniques and Libraries\n1. **Hugging Face Transformers**: A library providing implementations of transformer models, facilitating easy fine-tuning and deployment.\n2. **SpaCy**: An NLP library used for text preprocessing, tokenization, and NER.\n3. **NLTK (Natural Language Toolkit)**: A library for working with human language data, often used for tokenization, stemming, and other text processing tasks.\n4. **Pandas**: Used for handling datasets and transforming them into a format suitable for training models.\n\n### Training and Evaluation\n1. **Training Data**: The QNA dataset list is used as the training data, where each entry consists of a question (problem statement) and an answer (patch or solution).\n2. **Fine-Tuning**: The pre-trained models are fine-tuned on the specific QNA dataset to adapt them to the task of summarizing GitHub repositories.\n3. **Evaluation Metrics**: Metrics like BLEU, ROUGE, and METEOR are used to evaluate the quality of the generated summaries.\n\n### Summary\n\n- **ML Techniques**: Supervised learning, transfer learning\n- **NLP Techniques**: Tokenization, text preprocessing, NER, POS tagging, word embeddings\n- **LLMs and SLMs**: BERT, GPT, T5, BART\n- **Libraries**: Hugging Face Transformers, SpaCy, NLTK, Pandas\n- **Training and Evaluation**: Using the QNA dataset list for fine-tuning models and evaluating their performance\n","metadata":{}},{"cell_type":"markdown","source":"## 4.2. Let's Explore Some Sample ML Code snippets for this\nHere's a brief overview of each model along with some examples:\n\n### Gemma\n**Gemma** is an open language model developed by Google DeepMind. It's designed to be lightweight yet powerful, with models ranging from 2B to 27B parameters. Gemma is great for tasks like text generation, summarization, and more.\n\n**Example**: Generating a summary of a news article:\n```python\nfrom transformers import GemmaTokenizer, GemmaForCausalLM\n\ntokenizer = GemmaTokenizer.from_pretrained(\"google/gemma-2b\")\nmodel = GemmaForCausalLM.from_pretrained(\"google/gemma-2b\")\n\ninputs = tokenizer(\"The latest cricket match between India and Australia ended in a thrilling draw.\", return_tensors=\"pt\")\noutputs = model.generate(inputs[\"input_ids\"], max_length=50)\nprint(tokenizer.decode(outputs[0], skip_special_tokens=True))\n```\n\n### LLaMA\n**LLaMA** (Large Language Model Meta AI) is a series of open-source language models developed by Meta. These models are known for their efficiency and performance on various NLP tasks.\n\n**Example**: Generating a response to a user query:\n```python\nfrom transformers import LlamaTokenizer, LlamaForCausalLM\n\ntokenizer = LlamaTokenizer.from_pretrained(\"meta-llama/Llama-2-7B\")\nmodel = LlamaForCausalLM.from_pretrained(\"meta-llama/Llama-2-7B\")\n\ninputs = tokenizer(\"What is the capital of France?\", return_tensors=\"pt\")\noutputs = model.generate(inputs[\"input_ids\"], max_length=50)\nprint(tokenizer.decode(outputs[0], skip_special_tokens=True))\n```\n\n### BERT\n**BERT** (Bidirectional Encoder Representations from Transformers) is a transformer-based model developed by Google, primarily used for tasks like text classification, named entity recognition, and more.\n\n**Example**: Sentiment analysis on movie reviews:\n```python\nfrom transformers import BertTokenizer, BertForSequenceClassification\nfrom transformers import Trainer, TrainingArguments\n\ntokenizer = BertTokenizer.from_pretrained(\"bert-base-uncased\")\nmodel = BertForSequenceClassification.from_pretrained(\"bert-base-uncased\")\n\ntrain_dataset = ...  # Load your training dataset\ntest_dataset = ...  # Load your test dataset\n\ntraining_args = TrainingArguments(\n    output_dir=\"./results\",\n    num_train_epochs=3,\n    per_device_train_batch_size=16,\n    per_device_eval_batch_size=64,\n    evaluation_strategy=\"epoch\",\n    save_steps=10_000,\n    save_total_limit=2,\n    logging_dir=\"./logs\",\n)\n\ntrainer = Trainer(\n    model=model,\n    args=training_args,\n    train_dataset=train_dataset,\n    eval_dataset=test_dataset,\n)\n\ntrainer.train()\n```\n\n### Google Gemini Flash 2.0\n**Google Gemini Flash 2.0** is a hypothetical model, as there isn't an official release by that name yet. However, if it were to exist, it would likely be an advanced version of Google's Gemini models, designed for even faster and more efficient processing.\n\n**Example**: (Hypothetical) Real-time translation:\n```python\nfrom google_gemini_flash_2_0 import GeminiFlashTranslator\n\ntranslator = GeminiFlashTranslator.from_pretrained(\"google/gemini-flash-2.0\")\ntranslated_text = translator.translate(\"Hello, how are you?\", target_language=\"Spanish\")\nprint(translated_text)\n```\n\n","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}