Skip to content

About

Given a tab delimited list of news article ids and titles, use vector similarity and DBSCAN clustering to group similar articles and label the groups using keywords extracted using TF-IDF.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

15 Commits

Folders and files

Repository files navigation

Kitten shred img

Movie Kitten shred

Vibe coding process

  1. Create a prompt to create a prompt. E.g. Create a prompt for an LLM that instructs it to do this ... Prompt LLMs on LM Arena in side by side battle mode to create a prompt that will do what I want. In this case it was claude-opus-4-5-20251101-thinking-32k prompt.

  2. Use the generated Claude Sonnet 4.5 prompt as the prompt to generate the actual code. Which produced Claude artifact.

  3. Vibe fix code by feeding errors. This needed 2 main fixes, both to categorizer.py. These were fixed using Kimi K2.

  4. imports were missing

import logging
import numpy as np
import pandas as pd
from typing import Dict
from typing import List
from sklearn.cluster import DBSCAN
from sklearn.feature_extraction.text import TfidfVectorizer
  1. Convert numpy ints to Python ints in format_results()

This is different than the one in the Claude Artifact link

def format_results(
    categories: Dict[int, Dict],
    articles: pd.DataFrame,
    cluster_labels: np.ndarray
) -> Dict:
    """
    Format clustering results for output.
    
    Args:
        categories: Dictionary of category information
        articles: DataFrame with article data
        cluster_labels: Array of cluster assignments
        
    Returns:
        Formatted results dictionary
    """
    results = {
        'categories': {},
        'uncategorized': []
    }
    
    # Add categorized articles
    for cat_id, cat_info in categories.items():
        article_list = []
        for idx in cat_info['article_indices']:
            article_list.append({
                'article_id': str(articles.iloc[idx]['article_id']),
                'title': articles.iloc[idx]['title']
            })
        
        # Convert cat_id to int() to ensure it's a Python int, not numpy.int64
        display_id = int(cat_id) + 1  # 1-indexed for display
        
        results['categories'][display_id] = {
            'category_id': display_id,
            'category_name': cat_info['category_name'],
            'article_count': len(article_list),
            'articles': article_list
        }
    
    # Add uncategorized articles
    uncategorized_mask = cluster_labels == -1
    uncategorized_articles = articles[uncategorized_mask]
    
    for _, row in uncategorized_articles.iterrows():
        results['uncategorized'].append({
            'article_id': str(row['article_id']),
            'title': row['title']
        })
    
    return results

Claude Artifact embed code

<iframe src="https://claude.site/public/artifacts/7dcd1546-4e30-4344-bfc9-e92be72ea72e/embed" title="Claude Artifact" width="100%" height="600" frameborder="0" allow="clipboard-write" allowfullscreen></iframe>

and that's it, Folks! We now return you to your regularly scheduled repo ...

But no, for some reason it worked on my machine, but, like a good little geek I ran thru the install and found another error.

The error is caused by ChromaDB's delete_collection() throwing a NotFoundError (instead of ValueError) when the collection doesn't exist. The code only catches ValueError, so the exception isn't handled properly.

Look at git log for chroma_manager.py for changes.

About

Given a tab delimited list of news article ids and titles, use vector similarity and DBSCAN clustering to group similar articles and label the groups using keywords extracted using TF-IDF.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages