ysights.algorithms.paradox
Visibility Paradox Analysis
This module provides functions for analyzing the visibility paradox in social networks. The visibility paradox describes situations where content created by agents receives asymmetric visibility compared to content they see from their neighbors.
The module includes: - Visibility paradox detection and statistical significance testing - Comparison of user visibility vs. neighbor visibility - Population-size effects on the paradox - Null model generation for hypothesis testing
- Key Concepts:
The visibility paradox occurs when there’s an imbalance between: 1. How much of your neighbors’ content you see (inbound recommendations) 2. How much your neighbors see your content (outbound visibility)
This can lead to situations where most users feel their content is under-represented in their neighbors’ feeds, even though the aggregate statistics might suggest balance.
Example
Detecting the visibility paradox:
from ysights import YDataHandler
from ysights.algorithms.paradox import visibility_paradox, user_visibility_vs_neighbors
# Initialize data handler and extract network
ydh = YDataHandler('path/to/database.db')
network = ydh.social_network()
# Calculate visibility paradox with statistical testing
paradox_results = visibility_paradox(ydh, network, N=100)
print(f"Paradox score: {paradox_results['paradox_score']:.4f}")
print(f"Z-score: {paradox_results['z_score']:.4f}")
print(f"P-value: {paradox_results['p_value']:.4f}")
if paradox_results['p_value'] < 0.05:
print("Visibility paradox detected (statistically significant)")
# Compare user visibility with neighbor averages
user_vis, neighbor_vis = user_visibility_vs_neighbors(ydh, network)
import numpy as np
print(f"Average user visibility: {np.mean(user_vis):.2f}")
print(f"Average neighbor visibility: {np.mean(neighbor_vis):.2f}")
References
The visibility paradox is related to concepts from: - Friendship paradox (Feld, 1991) - Attention inequality in social networks - Filter bubble and echo chamber effects
For detailed mathematical formulation and theoretical background, see: docs/VISIBILITY_PARADOX.md - Complete mathematical description with formulas, assumptions, null model construction, and statistical testing methodology.
See also
visibility_paradox(): Main function for paradox detection
user_visibility_vs_neighbors(): Compare visibility metrics
visibility_paradox_population_size_null(): Analyze population size effects
Functions
|
Calculate the visibility for each user in the graph and the average of its neighbors' visibilities. |
|
Calculate the visibility paradox metric for a given YDataHandler and graph. |
Calculate the statistical significance of the visibility paradox per degree class. |
|
Calculate the visibility paradox metric for a given YDataHandler and graph, considering the population size. |
|
|
Calculate the visibility paradox over time with user-defined temporal granularity. |
- ysights.algorithms.paradox.user_visibility_vs_neighbors(YDH, g, node_ids=False)[source]
Calculate the visibility for each user in the graph and the average of its neighbors’ visibilities.
- Parameters:
YDH (
YDataHandler)g
- Returns:
- ysights.algorithms.paradox.visibility_paradox(YDH, g, N=100)[source]
Calculate the visibility paradox metric for a given YDataHandler and graph.
- Parameters:
YDH (
YDataHandler) – YDataHandler, the data handler containing the YSocial simulation datag – networkx.Graph, the social network graph
N – int, number of null models to generate for statistical testing
- Returns:
- ysights.algorithms.paradox.visibility_paradox_population_size_null(YDH, g, N=10, subject_to_rec=[0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9])[source]
Calculate the visibility paradox metric for a given YDataHandler and graph, considering the population size.
- Parameters:
YDH (
YDataHandler) – YDataHandler, the data handler containing the YSocial simulation datag – networkx.Graph, the social network graph
N – int, number of null models to generate for statistical testing
x – float, fraction of users to randomize (between 0 and 1)
- Returns:
- ysights.algorithms.paradox.visibility_paradox_per_degree_class(YDH, g, N=100, bins=None, num_bins=10)[source]
Calculate the statistical significance of the visibility paradox per degree class.
This function computes the paradox score and its statistical significance for each degree class (bin) of nodes. By default, it uses linear binning but allows users to specify custom bin edges.
- Parameters:
YDH (
YDataHandler) – YDataHandler, the data handler containing the YSocial simulation datag – networkx.Graph, the social network graph
N – int, number of null models to generate for statistical testing
bins – array-like, optional bin edges for degree binning. If None, linear bins are created.
num_bins – int, number of bins to create if bins is None (default: 10)
- Returns:
dict with keys: - ‘bin_edges’: array of bin edges - ‘bin_centers’: array of bin centers for plotting - ‘paradox_scores’: average paradox score per bin - ‘z_scores’: z-score per bin (statistical significance) - ‘p_values’: p-value per bin - ‘bin_counts’: number of nodes in each bin
Example
>>> from ysights import YDataHandler >>> from ysights.algorithms.paradox import visibility_paradox_per_degree_class >>> >>> ydh = YDataHandler('path/to/database.db') >>> network = ydh.social_network() >>> >>> # Using default linear binning >>> results = visibility_paradox_per_degree_class(ydh, network, N=100, num_bins=10) >>> >>> # Using custom bin edges >>> custom_bins = [0, 5, 10, 20, 50, 100] >>> results = visibility_paradox_per_degree_class(ydh, network, N=100, bins=custom_bins)
- ysights.algorithms.paradox.visibility_paradox_temporal(YDH, g, temporal_granularity=(1, 0), N=100)[source]
Calculate the visibility paradox over time with user-defined temporal granularity.
This function tracks how the visibility paradox evolves during the simulation by computing the paradox score at regular time intervals using INCREMENTAL/CUMULATIVE data. Each time point uses all data from the start of the simulation up to that time point, showing how the paradox strengthens or changes as more data accumulates.
- Parameters:
YDH (
YDataHandler) – YDataHandler, the data handler containing the YSocial simulation datag – networkx.Graph, the social network graph (can be full network or time-specific)
temporal_granularity – tuple of (days, hours) defining the time intervals e.g., (1, 0) = compute every 1 day, (0, 12) = every 12 hours, (1, 2) = every 26 hours (1 day + 2 hours)
N – int, number of null models to generate for statistical testing per time point
- Returns:
dict with keys: - ‘time_points’: list of (day, hour, round_id) tuples marking each cumulative endpoint - ‘paradox_scores’: array of paradox scores over time (cumulative) - ‘z_scores’: array of z-scores over time - ‘p_values’: array of p-values over time - ‘temporal_granularity’: the temporal granularity used (days, hours)
Note
The computation is INCREMENTAL - each time point includes all data from the simulation start up to that point. For example, with daily granularity: - Day 1: All data from start to day 1 - Day 2: All data from start to day 2 - Day 3: All data from start to day 3 This shows how the paradox evolves as the network and content accumulate.
Example
>>> from ysights import YDataHandler >>> from ysights.algorithms.paradox import visibility_paradox_temporal >>> >>> ydh = YDataHandler('path/to/database.db') >>> network = ydh.social_network() >>> >>> # Compute paradox every day (incrementally) >>> results = visibility_paradox_temporal(ydh, network, temporal_granularity=(1, 0), N=50) >>> >>> # Compute paradox every 12 hours (incrementally) >>> results = visibility_paradox_temporal(ydh, network, temporal_granularity=(0, 12), N=50) >>> >>> # Compute paradox every 26 hours (incrementally) >>> results = visibility_paradox_temporal(ydh, network, temporal_granularity=(1, 2), N=50)