Page MenuHomePhorge

No OneTemporary

Size
130 KB
Referenced Files
None
Subscribers
None
diff --git a/changelog.d/retention-mode.add b/changelog.d/retention-mode.add
index 898acb401..3e5625fb8 100644
--- a/changelog.d/retention-mode.add
+++ b/changelog.d/retention-mode.add
@@ -1 +1 @@
-Add a bounded retention mode (`config :pleroma, :retention`): a cron worker walks old activities in id order and evicts remote threads nobody local interacted with, in bounded batches, so the database plateaus instead of growing forever. Adds a `retention_cursors` table. `prune_objects --keep-threads` now also keeps threads that are addressed to a local user, still have a notification for one, or contain a reported post.
+Add a bounded retention mode (`config :pleroma, :retention`): a cron worker walks old activities in id order and evicts remote threads nobody local interacted with, in bounded batches, so the database plateaus instead of growing forever. Adds a `retention_cursors` table. `prune_objects --keep-threads` now also keeps threads that are addressed to a local user, still have a notification for one, contain a post quoted by a local post, or contain a reported post.
diff --git a/docs/configuration/cheatsheet.md b/docs/configuration/cheatsheet.md
index b299ef245..1bae9f966 100644
--- a/docs/configuration/cheatsheet.md
+++ b/docs/configuration/cheatsheet.md
@@ -1,1263 +1,1260 @@
# Configuration Cheat Sheet
This is a cheat sheet for Pleroma configuration file, any setting possible to configure should be listed here.
For OTP installations the configuration is typically stored in `/etc/pleroma/config.exs`.
For from source installations Pleroma configuration works by first importing the base config `config/config.exs`, then overriding it by the environment config `config/$MIX_ENV.exs` and then overriding it by user config `config/$MIX_ENV.secret.exs`. In from source installations you should always make the changes to the user config and NEVER to the base config to avoid breakages and merge conflicts. So for production you change/add configuration to `config/prod.secret.exs`.
To add configuration to your config file, you can copy it from the base config. The latest version of it can be viewed [here](https://git.pleroma.social/pleroma/pleroma/blob/develop/config/config.exs). You can also use this file if you don't know how an option is supposed to be formatted.
## :shout
* `enabled` - Enables the backend Shoutbox chat feature. Defaults to `true`.
* `limit` - Shout character limit. Defaults to `5_000`
## :instance
* `name`: The instance’s name.
* `email`: Email used to reach an Administrator/Moderator of the instance.
* `notify_email`: Email used for notifications.
* `description`: The instance’s description, can be seen in nodeinfo and ``/api/v1/instance``.
* `short_description`: Shorter version of instance description, can be seen on ``/api/v1/instance``.
* `limit`: Posts character limit (CW/Subject included in the counter).
* `description_limit`: The character limit for image descriptions.
* `remote_limit`: Hard character limit beyond which remote posts will be dropped.
* `upload_limit`: File size limit of uploads (except for avatar, background, banner).
* `avatar_upload_limit`: File size limit of user’s profile avatars.
* `background_upload_limit`: File size limit of user’s profile backgrounds.
* `banner_upload_limit`: File size limit of user’s profile banners.
* `poll_limits`: A map with poll limits for **local** polls.
* `max_options`: Maximum number of options.
* `max_option_chars`: Maximum number of characters per option.
* `min_expiration`: Minimum expiration time (in seconds).
* `max_expiration`: Maximum expiration time (in seconds).
* `registrations_open`: Enable registrations for anyone, invitations can be enabled when false.
* `invites_enabled`: Enable user invitations for admins (depends on `registrations_open: false`).
* `account_activation_required`: Require users to confirm their emails before signing in.
* `account_approval_required`: Require users to be manually approved by an admin before signing in.
* `federating`: Enable federation with other instances.
* `federation_incoming_replies_max_depth`: Max. depth of reply-to activities fetching on incoming federation, to prevent out-of-memory situations while fetching very long threads. If set to `nil`, threads of any depth will be fetched. Lower this value if you experience out-of-memory crashes.
* `federation_reachability_timeout_days`: Timeout (in days) of each external federation target being unreachable prior to pausing federating to it.
* `allow_relay`: Permits remote instances to subscribe to all public posts of your instance. This may increase the visibility of your instance.
* `public`: Makes the client API in authenticated mode-only except for user-profiles. Useful for disabling the Local Timeline and The Whole Known Network. Note that there is a dependent setting restricting or allowing unauthenticated access to specific resources, see `restrict_unauthenticated` for more details.
* `quarantined_instances`: ActivityPub instances where private (DMs, followers-only) activities will not be send.
* `rejected_instances`: ActivityPub instances to reject requests from if authorized_fetch_mode is enabled.
* `allowed_post_formats`: MIME-type list of formats allowed to be posted (transformed into HTML).
* `extended_nickname_format`: Set to `true` to use extended local nicknames format (allows underscores/dashes). This will break federation with
older software for theses nicknames.
* `max_pinned_statuses`: The maximum number of pinned statuses. `0` will disable the feature.
* `autofollowed_nicknames`: Set to nicknames of (local) users that every new user should automatically follow.
* `autofollowing_nicknames`: Set to nicknames of (local) users that automatically follows every newly registered user.
* `attachment_links`: Set to true to enable automatically adding attachment link text to statuses.
* `max_report_comment_size`: The maximum size of the report comment (Default: `1000`).
* `report_strip_status`: Strip associated statuses in reports to ids when closed/resolved, otherwise keep a copy.
* `safe_dm_mentions`: If set to true, only mentions at the beginning of a post will be used to address people in direct messages. This is to prevent accidental mentioning of people when talking about them (e.g. "@friend hey i really don't like @enemy"). Default: `false`.
* `healthcheck`: If set to true, system data will be shown on ``/api/v1/pleroma/healthcheck``.
* `remote_post_retention_days`: The default amount of days to retain remote posts when pruning the database.
* `user_bio_length`: A user bio maximum length (default: `5000`).
* `user_name_length`: A user name maximum length (default: `100`).
* `skip_thread_containment`: Skip filter out broken threads. The default is `false`.
* `limit_to_local_content`: Limit unauthenticated users to search for local statutes and users only. Possible values: `:unauthenticated`, `:all` and `false`. The default is `:unauthenticated`.
* `max_account_fields`: The maximum number of custom fields in the user profile (default: `10`).
* `max_remote_account_fields`: The maximum number of custom fields in the remote user profile (default: `20`).
* `account_field_name_length`: An account field name maximum length (default: `512`).
* `account_field_value_length`: An account field value maximum length (default: `2048`).
* `registration_reason_length`: Maximum registration reason length (default: `500`).
* `external_user_synchronization`: Enabling following/followers counters synchronization for external users.
* `cleanup_attachments`: Remove attachments along with statuses. Does not affect duplicate files and attachments without status. Enabling this will increase load to database when deleting statuses on larger instances.
* `show_reactions`: Let favourites and emoji reactions be viewed through the API (default: `true`).
* `password_reset_token_validity`: The time after which reset tokens aren't accepted anymore, in seconds (default: one day).
* `admin_privileges`: A list of privileges an admin has (e.g. delete messages, manage reports...)
* Possible values are:
* `:users_read`
* Allows admins to fetch users through the admin API.
* `:users_manage_invites`
* Allows admins to manage invites. This includes sending, resending, revoking and approving invites.
* `:users_manage_activation_state`
* Allows admins to activate and deactivate accounts. This also allows them to see deactivated users through the Mastodon API.
* `:users_manage_tags`
* Allows admins to set and remove tags for users. This can be useful in combination with MRF policies, such as `Pleroma.Web.ActivityPub.MRF.TagPolicy`.
* `:users_manage_credentials`
* Allows admins to trigger a password reset and set new credentials for an user.
* `:users_delete`
* Allows admins to delete accounts. Note that deleting an account is actually deactivating it and removing all data like posts, profile information, etc.
* `:messages_read`
* Allows admins to read messages through the admin API, including non-public posts and chats.
* `:messages_delete`
* Allows admins to delete messages from other users.
* `:instances_delete,`
* Allows admins to remove a whole remote instance from your instance. This will delete all users and messages from that remote instance.
* `:reports_manage_reports`
* Allows admins to see and manage reports.
* `:moderation_log_read,`
* Allows admins to read the entries in the moderation log.
* `:emoji_manage_emoji`
* Allows admins to manage custom emoji on the instance.
* `:statistics_read,`
* Allows admins to see some simple statistics about the instance.
* `moderator_privileges`: A list of privileges a moderator has (e.g. delete messages, manage reports...)
* Possible values are the same as for `admin_privileges`
## :features
* `improved_hashtag_timeline`: Setting to force toggle / force disable improved hashtags timeline. `:enabled` forces hashtags to be fetched from `hashtags` table for hashtags timeline. `:disabled` forces object-embedded hashtags to be used (slower). Keep it `:auto` for automatic behaviour (it is auto-set to `:enabled` [unless overridden] when HashtagsTableMigrator completes).
## Background migrations
* `populate_hashtags_table/sleep_interval_ms`: Sleep interval between each chunk of processed records in order to decrease the load on the system (defaults to 0 and should be keep default on most instances).
* `populate_hashtags_table/fault_rate_allowance`: Max rate of failed objects to actually processed objects in order to enable the feature (any value from 0.0 which tolerates no errors to 1.0 which will enable the feature even if hashtags transfer failed for all records).
## Welcome
* `direct_message`: - welcome message sent as a direct message.
* `enabled`: Enables the send a direct message to a newly registered user. Defaults to `false`.
* `sender_nickname`: The nickname of the local user that sends the welcome message.
* `message`: A message that will be send to a newly registered users as a direct message.
* `chat_message`: - welcome message sent as a chat message.
* `enabled`: Enables the send a chat message to a newly registered user. Defaults to `false`.
* `sender_nickname`: The nickname of the local user that sends the welcome message.
* `message`: A message that will be send to a newly registered users as a chat message.
* `email`: - welcome message sent as a email.
* `enabled`: Enables the send a welcome email to a newly registered user. Defaults to `false`.
* `sender`: The email address or tuple with `{nickname, email}` that will use as sender to the welcome email.
* `subject`: A subject of welcome email.
* `html`: A html that will be send to a newly registered users as a email.
* `text`: A text that will be send to a newly registered users as a email.
Example:
```elixir
config :pleroma, :welcome,
direct_message: [
enabled: true,
sender_nickname: "lain",
message: "Hi! Welcome on board!"
],
email: [
enabled: true,
sender: {"Pleroma App", "welcome@pleroma.app"},
subject: "Welcome to <%= instance_name %>",
html: "Welcome to <%= instance_name %>",
text: "Welcome to <%= instance_name %>"
]
```
## Message rewrite facility
### :mrf
* `policies`: Message Rewrite Policy, either one or a list. Here are the ones available by default:
* `Pleroma.Web.ActivityPub.MRF.NoOpPolicy`: Doesn’t modify activities (default).
* `Pleroma.Web.ActivityPub.MRF.DropPolicy`: Drops all activities. It generally doesn’t makes sense to use in production.
* `Pleroma.Web.ActivityPub.MRF.SimplePolicy`: Restrict the visibility of activities from certains instances (See [`:mrf_simple`](#mrf_simple)).
* `Pleroma.Web.ActivityPub.MRF.TagPolicy`: Applies policies to individual users based on tags, which can be set using pleroma-fe/admin-fe/any other app that supports Pleroma Admin API. For example it allows marking posts from individual users nsfw (sensitive).
* `Pleroma.Web.ActivityPub.MRF.SubchainPolicy`: Selectively runs other MRF policies when messages match (See [`:mrf_subchain`](#mrf_subchain)).
* `Pleroma.Web.ActivityPub.MRF.RejectNonPublic`: Drops posts with non-public visibility settings (See [`:mrf_rejectnonpublic`](#mrf_rejectnonpublic)).
* `Pleroma.Web.ActivityPub.MRF.EnsureRePrepended`: Rewrites posts to ensure that replies to posts with subjects do not have an identical subject and instead begin with re:.
* `Pleroma.Web.ActivityPub.MRF.AntiLinkSpamPolicy`: Rejects posts from likely spambots by rejecting posts from new users that contain links.
* `Pleroma.Web.ActivityPub.MRF.MediaProxyWarmingPolicy`: Crawls attachments using their MediaProxy URLs so that the MediaProxy cache is primed.
* `Pleroma.Web.ActivityPub.MRF.MentionPolicy`: Drops posts mentioning configurable users. (See [`:mrf_mention`](#mrf_mention)).
* `Pleroma.Web.ActivityPub.MRF.VocabularyPolicy`: Restricts activities to a configured set of vocabulary. (See [`:mrf_vocabulary`](#mrf_vocabulary)).
* `Pleroma.Web.ActivityPub.MRF.ObjectAgePolicy`: Rejects or delists posts based on their age when received. (See [`:mrf_object_age`](#mrf_object_age)).
* `Pleroma.Web.ActivityPub.MRF.ActivityExpirationPolicy`: Sets a default expiration on all posts made by users of the local instance. Requires `Pleroma.Workers.PurgeExpiredActivity` to be enabled for processing the scheduled deletions.
* `Pleroma.Web.ActivityPub.MRF.ForceBotUnlistedPolicy`: Makes all bot posts to disappear from public timelines.
* `Pleroma.Web.ActivityPub.MRF.FollowBotPolicy`: Automatically follows newly discovered users from the specified bot account. Local accounts, locked accounts, and users with "#nobot" in their bio are respected and excluded from being followed.
* `Pleroma.Web.ActivityPub.MRF.AntiFollowbotPolicy`: Drops follow requests from followbots. Users can still allow bots to follow them by first following the bot.
* `Pleroma.Web.ActivityPub.MRF.KeywordPolicy`: Rejects or removes from the federated timeline or replaces keywords. (See [`:mrf_keyword`](#mrf_keyword)).
* `Pleroma.Web.ActivityPub.MRF.ForceMentionsInContent`: Forces every mentioned user to be reflected in the post content.
* `Pleroma.Web.ActivityPub.MRF.InlineQuotePolicy`: Forces quote post URLs to be reflected in the message content inline.
* `Pleroma.Web.ActivityPub.MRF.QuoteToLinkTagPolicy`: Force a Link tag for posts quoting another post. (may break outgoing federation of quote posts with older Pleroma versions).
* `Pleroma.Web.ActivityPub.MRF.ForceMention`: Forces posts to include a mention of the author of parent post or the author of quoted post.
* `transparency`: Make the content of your Message Rewrite Facility settings public (via nodeinfo).
* `transparency_exclusions`: Exclude specific instance names from MRF transparency. The use of the exclusions feature will be disclosed in nodeinfo as a boolean value.
## Federation
### MRF policies
!!! note
Configuring MRF policies is not enough for them to take effect. You have to enable them by specifying their module in `policies` under [:mrf](#mrf) section.
#### :mrf_simple
* `media_removal`: List of instances to strip media attachments from and the reason for doing so.
* `media_nsfw`: List of instances to tag all media as NSFW (sensitive) from and the reason for doing so.
* `federated_timeline_removal`: List of instances to remove from the Federated Timeline (aka The Whole Known Network) and the reason for doing so.
* `reject`: List of instances to reject activities (except deletes) from and the reason for doing so.
* `accept`: List of instances to only accept activities (except deletes) from and the reason for doing so.
* `followers_only`: Force posts from the given instances to be visible by followers only and the reason for doing so.
* `report_removal`: List of instances to reject reports from and the reason for doing so.
* `avatar_removal`: List of instances to strip avatars from and the reason for doing so.
* `banner_removal`: List of instances to strip banners from and the reason for doing so.
* `reject_deletes`: List of instances to reject deletions from and the reason for doing so.
#### :mrf_subchain
This policy processes messages through an alternate pipeline when a given message matches certain criteria.
All criteria are configured as a map of regular expressions to lists of policy modules.
* `match_actor`: Matches a series of regular expressions against the actor field.
Example:
```elixir
config :pleroma, :mrf_subchain,
match_actor: %{
~r/https:\/\/example.com/s => [Pleroma.Web.ActivityPub.MRF.DropPolicy]
}
```
#### :mrf_rejectnonpublic
* `allow_followersonly`: whether to allow followers-only posts.
* `allow_direct`: whether to allow direct messages.
#### :mrf_hellthread
* `delist_threshold`: Number of mentioned users after which the message gets delisted (the message can still be seen, but it will not show up in public timelines and mentioned users won't get notifications about it). Set to 0 to disable.
* `reject_threshold`: Number of mentioned users after which the messaged gets rejected. Set to 0 to disable.
#### :mrf_keyword
* `reject`: A list of patterns which result in message being rejected, each pattern can be a string or a [regular expression](https://hexdocs.pm/elixir/Regex.html).
* `federated_timeline_removal`: A list of patterns which result in message being removed from federated timelines (a.k.a unlisted), each pattern can be a string or a [regular expression](https://hexdocs.pm/elixir/Regex.html).
* `replace`: A list of tuples containing `{pattern, replacement}`, `pattern` can be a string or a [regular expression](https://hexdocs.pm/elixir/Regex.html).
#### :mrf_mention
* `actors`: A list of actors, for which to drop any posts mentioning.
#### :mrf_vocabulary
* `accept`: A list of ActivityStreams terms to accept. If empty, all supported messages are accepted.
* `reject`: A list of ActivityStreams terms to reject. If empty, no messages are rejected.
#### :mrf_user_allowlist
The keys in this section are the domain names that the policy should apply to.
Each key should be assigned a list of users that should be allowed through by
their ActivityPub ID.
An example:
```elixir
config :pleroma, :mrf_user_allowlist, %{
"example.org" => ["https://example.org/users/admin"]
}
```
#### :mrf_object_age
* `threshold`: Required time offset (in seconds) compared to your server clock of an incoming post before actions are taken.
e.g., A value of 900 results in any post with a timestamp older than 15 minutes will be acted upon.
* `actions`: A list of actions to apply to the post:
* `:delist` removes the post from public timelines
* `:strip_followers` removes followers from the ActivityPub recipient list, ensuring they won't be delivered to home timelines, additionally for followers-only it degrades to a direct message
* `:reject` rejects the message entirely
#### :mrf_steal_emoji
* `hosts`: List of hosts to steal emojis from
* `rejected_shortcodes`: Regex-list of shortcodes to reject
* `size_limit`: File size limit (in bytes), checked before an emoji is saved to the disk
#### :mrf_activity_expiration
* `days`: Default global expiration time for all local Create activities (in days)
#### :mrf_hashtag
* `sensitive`: List of hashtags to mark activities as sensitive (default: `nsfw`)
* `federated_timeline_removal`: List of hashtags to remove activities from the federated timeline (aka TWNK)
* `reject`: List of hashtags to reject activities from
Notes:
- The hashtags in the configuration do not have a leading `#`.
- This MRF Policy is always enabled, if you want to disable it you have to set empty lists
#### :mrf_follow_bot
* `follower_nickname`: The name of the bot account to use for following newly discovered users. Using `followbot` or similar is strongly suggested.
#### :mrf_emoji
* `remove_url`: A list of patterns which result in emoji whose URL matches being removed from the message. This will apply to statuses, emoji reactions, and user profiles. Each pattern can be a string or a [regular expression](https://hexdocs.pm/elixir/Regex.html).
* `remove_shortcode`: A list of patterns which result in emoji whose shortcode matches being removed from the message. This will apply to statuses, emoji reactions, and user profiles. Each pattern can be a string or a [regular expression](https://hexdocs.pm/elixir/Regex.html).
* `federated_timeline_removal_url`: A list of patterns which result in message with emojis whose URLs match being removed from federated timelines (a.k.a unlisted). This will apply only to statuses. Each pattern can be a string or a [regular expression](https://hexdocs.pm/elixir/Regex.html).
* `federated_timeline_removal_shortcode`: A list of patterns which result in message with emojis whose shortcodes match being removed from federated timelines (a.k.a unlisted). This will apply only to statuses. Each pattern can be a string or a [regular expression](https://hexdocs.pm/elixir/Regex.html).
#### :mrf_inline_quote
* `template`: The template to append to the post. `{url}` will be replaced with the actual link to the quoted post. Default: `<bdi>RT:</bdi> {url}`
#### :mrf_force_mention
* `mention_parent`: Whether to append mention of parent post author
* `mention_quoted`: Whether to append mention of parent quoted author
### :activitypub
* `unfollow_blocked`: Whether blocks result in people getting unfollowed
* `outgoing_blocks`: Whether to federate blocks to other instances
* `blockers_visible`: Whether a user can see the posts of users who blocked them
* `deny_follow_blocked`: Whether to disallow following an account that has blocked the user in question
* `sign_object_fetches`: Sign object fetches with HTTP signatures
* `authorized_fetch_mode`: Require HTTP signatures for AP fetches
* `authorized_fetch_mode_exceptions`: List of IPs (CIDR format accepted) to exempt from HTTP Signatures requirement (for example to allow debugging, you shouldn't otherwise need this)
## Pleroma.User
* `restricted_nicknames`: List of nicknames users may not register with.
* `email_blacklist`: List of email domains users may not register with.
## Pleroma.ScheduledActivity
* `daily_user_limit`: the number of scheduled activities a user is allowed to create in a single day (Default: `25`)
* `total_user_limit`: the number of scheduled activities a user is allowed to create in total (Default: `300`)
* `enabled`: whether scheduled activities are sent to the job queue to be executed
### :frontend_configurations
This can be used to configure a keyword list that keeps the configuration data for any kind of frontend. By default, settings for `pleroma_fe` are configured. You can find the documentation for `pleroma_fe` configuration into [Pleroma-FE configuration and customization for instance administrators](/frontend/CONFIGURATION/#options).
Frontends can access these settings at `/api/v1/pleroma/frontend_configurations`
To add your own configuration for PleromaFE, use it like this:
```elixir
config :pleroma, :frontend_configurations,
pleroma_fe: %{
theme: "pleroma-dark",
# ... see /priv/static/static/config.json for the available keys.
}
```
These settings **need to be complete**, they will override the defaults.
### :static_fe
Render profiles and posts using server-generated HTML that is viewable without using JavaScript.
Available options:
* `enabled` - Enables the rendering of static HTML. Defaults to `false`.
### :assets
This section configures assets to be used with various frontends. Currently the only option
relates to mascots on the mastodon frontend
* `mascots`: KeywordList of mascots, each element __MUST__ contain both a `url` and a
`mime_type` key.
* `default_mascot`: An element from `mascots` - This will be used as the default mascot
on MastoFE (default: `:pleroma_fox_tan`).
### :manifest
This section describe PWA manifest instance-specific values. Currently this option relate only for MastoFE.
* `icons`: Describe the icons of the app, this a list of maps describing icons in the same way as the
[spec](https://www.w3.org/TR/appmanifest/#imageresource-and-its-members) describes it.
Example:
```elixir
config :pleroma, :manifest,
icons: [
%{
src: "/static/logo.png"
},
%{
src: "/static/icon.png",
type: "image/png"
},
%{
src: "/static/icon.ico",
sizes: "72x72 96x96 128x128 256x256"
}
]
```
* `theme_color`: Describe the theme color of the app. (Example: `"#282c37"`, `"rebeccapurple"`).
* `background_color`: Describe the background color of the app. (Example: `"#191b22"`, `"aliceblue"`).
## :emoji
* `shortcode_globs`: Location of custom emoji files. `*` can be used as a wildcard. Example `["/emoji/custom/**/*.png"]`
* `pack_extensions`: A list of file extensions for emojis, when no emoji.txt for a pack is present. Example `[".png", ".gif"]`
* `groups`: Emojis are ordered in groups (tags). This is an array of key-value pairs where the key is the groupname and the value the location or array of locations. `*` can be used as a wildcard. Example `[Custom: ["/emoji/*.png", "/emoji/custom/*.png"]]`
* `default_manifest`: Location of the JSON-manifest. This manifest contains information about the emoji-packs you can download. Currently only one manifest can be added (no arrays).
* `shared_pack_cache_seconds_per_file`: When an emoji pack is shared, the archive is created and cached in
memory for this amount of seconds multiplied by the number of files.
## :media_proxy
* `enabled`: Enables proxying of remote media to the instance’s proxy
* `base_url`: The base URL to access a user-uploaded file. Useful when you want to proxy the media files via another host/CDN fronts.
* `proxy_opts`: All options defined in `Pleroma.ReverseProxy` documentation, defaults to `[max_body_length: (25*1_048_576)]`.
* `whitelist`: List of hosts with scheme to bypass the mediaproxy (e.g. `https://example.com`)
* `invalidation`: options for remove media from cache after delete object:
* `enabled`: Enables purge cache
* `provider`: Which one of the [purge cache strategy](#purge-cache-strategy) to use.
## :media_preview_proxy
* `enabled`: Enables proxying of remote media preview to the instance’s proxy. Requires enabled media proxy (`media_proxy/enabled`).
* `thumbnail_max_width`: Max width of preview thumbnail for images (video preview always has original dimensions).
* `thumbnail_max_height`: Max height of preview thumbnail for images (video preview always has original dimensions).
* `image_quality`: Quality of the output. Ranges from 0 (min quality) to 100 (max quality).
* `min_content_length`: Min content length to perform preview, in bytes. If greater than 0, media smaller in size will be served as is, without thumbnailing.
### Purge cache strategy
#### Pleroma.Web.MediaProxy.Invalidation.Script
This strategy allow perform external shell script to purge cache.
Urls of attachments are passed to the script as arguments.
* `script_path`: Path to the external script.
* `url_format`: Set to `:htcacheclean` if using Apache's htcacheclean utility.
Example:
```elixir
config :pleroma, Pleroma.Web.MediaProxy.Invalidation.Script,
script_path: "./installation/nginx-cache-purge.example"
```
#### Pleroma.Web.MediaProxy.Invalidation.Http
This strategy allow perform custom http request to purge cache.
* `method`: http method. default is `purge`
* `headers`: http headers.
* `options`: request options.
Example:
```elixir
config :pleroma, Pleroma.Web.MediaProxy.Invalidation.Http,
method: :purge,
headers: [],
options: []
```
## Link previews
### Pleroma.Web.Metadata (provider)
* `providers`: a list of metadata providers to enable. Providers available:
* `Pleroma.Web.Metadata.Providers.OpenGraph`
* `Pleroma.Web.Metadata.Providers.TwitterCard`
* `unfurl_nsfw`: If set to `true` nsfw attachments will be shown in previews.
### :rich_media (consumer)
* `enabled`: if enabled the instance will parse metadata from attached links to generate link previews.
* `ignore_hosts`: list of hosts which will be ignored by the metadata parser. For example `["accounts.google.com", "xss.website"]`, defaults to `[]`.
* `ignore_tld`: list TLDs (top-level domains) which will ignore for parse metadata. default is ["local", "localdomain", "lan"].
* `parsers`: list of Rich Media parsers.
* `timeout`: Amount of milliseconds after which the HTTP request is forcibly terminated.
## HTTP server
### Pleroma.Web.Endpoint
!!! note
`Phoenix` endpoint configuration, all configuration options can be viewed [here](https://hexdocs.pm/phoenix/Phoenix.Endpoint.html#module-dynamic-configuration), only common options are listed here.
* `http` - a list containing http protocol configuration, all configuration options can be viewed [here](https://hexdocs.pm/plug_cowboy/Plug.Cowboy.html#module-options), only common options are listed here. For deployment using docker, you need to set this to `[ip: {0,0,0,0}, port: 4000]` to make pleroma accessible from other containers (such as your nginx server).
- `ip` - a tuple consisting of 4 integers
- `port`
* `url` - a list containing the configuration for generating urls, accepts
- `host` - the host without the scheme and a post (e.g `example.com`, not `https://example.com:2020`)
- `scheme` - e.g `http`, `https`
- `port`
- `path`
* `extra_cookie_attrs` - a list of `Key=Value` strings to be added as non-standard cookie attributes. Defaults to `["SameSite=Lax"]`. See the [SameSite article](https://www.owasp.org/index.php/SameSite) on OWASP for more info.
Example:
```elixir
config :pleroma, Pleroma.Web.Endpoint,
url: [host: "example.com", port: 2020, scheme: "https"],
http: [
port: 8080,
ip: {127, 0, 0, 1}
]
```
This will make Pleroma listen on `127.0.0.1` port `8080` and generate urls starting with `https://example.com:2020`
### :http_security
* ``enabled``: Whether the managed content security policy is enabled.
* ``sts``: Whether to additionally send a `Strict-Transport-Security` header.
* ``sts_max_age``: The maximum age for the `Strict-Transport-Security` header if sent.
* ``ct_max_age``: The maximum age for the `Expect-CT` header if sent.
* ``referrer_policy``: The referrer policy to use, either `"same-origin"` or `"no-referrer"`.
* ``report_uri``: Adds the specified url to `report-uri` and `report-to` group in CSP header.
* `allow_unsafe_eval`: Adds `wasm-unsafe-eval` to the CSP header. Needed for some non-essential frontend features like Flash emulation.
### Pleroma.Web.Plugs.RemoteIp
!!! warning
If your instance is not behind at least one reverse proxy, you should not enable this plug.
`Pleroma.Web.Plugs.RemoteIp` is a shim to call [`RemoteIp`](https://git.pleroma.social/pleroma/remote_ip) but with runtime configuration.
Available options:
* `enabled` - Enable/disable the plug. Defaults to `false`.
* `headers` - A list of strings naming the HTTP headers to use when deriving the true client IP address. Defaults to `["x-forwarded-for"]`.
* `proxies` - A list of upstream proxy IP subnets in CIDR notation from which we will parse the content of `headers`. Defaults to `[]`. IPv4 entries without a bitmask will be assumed to be /32 and IPv6 /128.
* `reserved` - A list of reserved IP subnets in CIDR notation which should be ignored if found in `headers`. Defaults to `["127.0.0.0/8", "::1/128", "fc00::/7", "10.0.0.0/8", "172.16.0.0/12", "192.168.0.0/16"]`.
### :rate_limit
!!! note
If your instance is behind a reverse proxy ensure [`Pleroma.Web.Plugs.RemoteIp`](#pleroma-plugs-remoteip) is enabled (it is enabled by default).
A keyword list of rate limiters where a key is a limiter name and value is the limiter configuration. The basic configuration is a tuple where:
* The first element: `scale` (Integer). The time scale in milliseconds.
* The second element: `limit` (Integer). How many requests to limit in the time scale provided.
It is also possible to have different limits for unauthenticated and authenticated users: the keyword value must be a list of two tuples where the first one is a config for unauthenticated users and the second one is for authenticated.
For example:
```elixir
config :pleroma, :rate_limit,
authentication: {60_000, 15},
search: [{1000, 10}, {1000, 30}]
```
Means that:
1. In 60 seconds, 15 authentication attempts can be performed from the same IP address.
2. In 1 second, 10 search requests can be performed from the same IP address by unauthenticated users, while authenticated users can perform 30 search requests per second.
Supported rate limiters:
* `:search` - Account/Status search.
* `:timeline` - Timeline requests (each timeline has it's own limiter).
* `:app_account_creation` - Account registration from the API.
* `:relations_actions` - Following/Unfollowing in general.
* `:relation_id_action` - Following/Unfollowing for a specific user.
* `:statuses_actions` - Status actions such as: (un)repeating, (un)favouriting, creating, deleting.
* `:status_id_action` - (un)Repeating/(un)Favouriting a particular status.
* `:authentication` - Authentication actions, i.e getting an OAuth token.
* `:password_reset` - Requesting password reset emails.
* `:account_confirmation_resend` - Requesting resending account confirmation emails.
* `:ap_routes` - Requesting statuses via ActivityPub.
### :web_cache_ttl
The expiration time for the web responses cache. Values should be in milliseconds or `nil` to disable expiration.
Available caches:
* `:activity_pub` - activity pub routes (except question activities). Defaults to `nil` (no expiration).
* `:activity_pub_question` - activity pub routes (question activities). Defaults to `30_000` (30 seconds).
## HTTP client
### :http
* `proxy_url`: an upstream proxy to fetch posts and/or media with, (default: `nil`)
* `send_user_agent`: should we include a user agent with HTTP requests? (default: `true`)
* `user_agent`: what user agent should we use? (default: `:default`), must be string or `:default`
* `adapter`: array of adapter options
### :hackney_pools
Advanced. Tweaks Hackney (http client) connections pools.
There's three pools used:
* `:federation` for the federation jobs.
You may want this pool max_connections to be at least equal to the number of federator jobs + retry queue jobs.
* `:media` for rich media, media proxy
* `:upload` for uploaded media (if using a remote uploader and `proxy_remote: true`)
For each pool, the options are:
* `max_connections` - how much connections a pool can hold
* `timeout` - retention duration for connections
### :connections_pool
*For `gun` adapter*
Settings for HTTP connection pool.
* `:connection_acquisition_wait` - Timeout to acquire a connection from pool.The total max time is this value multiplied by the number of retries.
* `connection_acquisition_retries` - Number of attempts to acquire the connection from the pool if it is overloaded. Each attempt is timed `:connection_acquisition_wait` apart.
* `:max_connections` - Maximum total number of connections in the pool. HTTP/1 origins may use multiple connections within this limit, while HTTP/2 connections are multiplexed.
* `:max_idle_time` - Time before an unused connection is closed.
* `:connect_timeout` - Timeout to connect to the host.
* `:reclaim_multiplier` - Multiplied by `:max_connections` this will be the maximum number of idle connections that will be reclaimed in case the pool is overloaded.
### :pools
*For `gun` adapter*
Settings for request pools. These pools are limited on top of `:connections_pool`.
There are four pools used:
* `:federation` for the federation jobs. You may want this pool's max_connections to be at least equal to the number of federator jobs + retry queue jobs.
* `:media` - for rich media, media proxy.
* `:upload` - for proxying media when a remote uploader is used and `proxy_remote: true`.
* `:default` - for other requests.
For each pool, the options are:
* `:size` - limit to how much requests can be concurrently executed.
* `:recv_timeout` - timeout while `gun` will wait for response
* `:max_waiting` - limit to how much requests can be waiting for others to finish, after this is reached, subsequent requests will be dropped.
## Captcha
### Pleroma.Captcha
* `enabled`: Whether the captcha should be shown on registration.
* `method`: The method/service to use for captcha.
* `seconds_valid`: The time in seconds for which the captcha is valid.
### Captcha providers
#### Pleroma.Captcha.Native
A built-in captcha provider. Enabled by default.
#### Pleroma.Captcha.Kocaptcha
Kocaptcha is a very simple captcha service with a single API endpoint,
the source code is here: [kocaptcha](https://github.com/koto-bank/kocaptcha). The default endpoint
`https://captcha.kotobank.ch` is hosted by the developer.
* `endpoint`: the Kocaptcha endpoint to use.
## Uploads
### Pleroma.Upload
* `uploader`: Which one of the [uploaders](#uploaders) to use.
* `filters`: List of [upload filters](#upload-filters) to use.
* `link_name`: When enabled Pleroma will add a `name` parameter to the url of the upload, for example `https://instance.tld/media/corndog.png?name=corndog.png`. This is needed to provide the correct filename in Content-Disposition headers when using filters like `Pleroma.Upload.Filter.Dedupe`
* `base_url`: The base URL to access a user-uploaded file. Useful when you want to host the media files via another domain or are using a 3rd party S3 provider.
* `proxy_remote`: If you're using a remote uploader, Pleroma will proxy media requests instead of redirecting to it.
* `proxy_opts`: Proxy options, see `Pleroma.ReverseProxy` documentation.
* `filename_display_max_length`: Set max length of a filename to display. 0 = no limit. Default: 30.
* `default_description`: Sets which default description an image has if none is set explicitly. Options: nil (default) - Don't set a default, :filename - use the filename of the file, a string (e.g. "attachment") - Use this string
!!! warning
`strip_exif` has been replaced by `Pleroma.Upload.Filter.Mogrify`.
### Uploaders
#### Pleroma.Uploaders.Local
* `uploads`: Which directory to store the user-uploads in, relative to pleroma’s working directory.
#### Pleroma.Uploaders.S3
Don't forget to configure [Ex AWS S3](#ex-aws-s3-settings)
* `bucket`: S3 bucket name.
* `bucket_namespace`: S3 bucket namespace.
* `truncated_namespace`: If you use S3 compatible service such as Digital Ocean Spaces or CDN, set folder name or "" etc.
* `streaming_enabled`: Enable streaming uploads, when enabled the file will be sent to the server in chunks as it's being read. This may be unsupported by some providers, try disabling this if you have upload problems.
#### Ex AWS S3 settings
* `access_key_id`: Access key ID
* `secret_access_key`: Secret access key
* `host`: S3 host
Example:
```elixir
config :ex_aws, :s3,
access_key_id: "xxxxxxxxxx",
secret_access_key: "yyyyyyyyyy",
host: "s3.eu-central-1.amazonaws.com"
```
#### Pleroma.Uploaders.IPFS
* `post_gateway_url`: URL with port of POST Gateway (unauthenticated)
* `get_gateway_url`: URL of public GET Gateway
Example:
```elixir
config :pleroma, Pleroma.Uploaders.IPFS,
post_gateway_url: "http://localhost:5001",
get_gateway_url: "http://{CID}.ipfs.mydomain.com"
```
### Upload filters
#### Pleroma.Upload.Filter.AnonymizeFilename
This filter replaces the filename (not the path) of an upload. For complete obfuscation, add
`Pleroma.Upload.Filter.Dedupe` before AnonymizeFilename.
* `text`: Text to replace filenames in links. If empty, `{random}.extension` will be used. You can get the original filename extension by using `{extension}`, for example `custom-file-name.{extension}`.
#### Pleroma.Upload.Filter.Dedupe
No specific configuration.
#### Pleroma.Upload.Filter.Exiftool.StripLocation
This filter only strips the GPS and location metadata with Exiftool leaving color profiles and attributes intact.
No specific configuration.
#### Pleroma.Upload.Filter.Exiftool.ReadDescription
This filter reads the ImageDescription and iptc:Caption-Abstract fields with Exiftool so clients can prefill the media description field.
No specific configuration.
#### Pleroma.Upload.Filter.OnlyMedia
This filter rejects uploads that are not identified with Content-Type matching audio/\*, image/\*, or video/\*
No specific configuration.
#### Pleroma.Upload.Filter.Mogrify
* `args`: List of actions for the `mogrify` command like `"strip"` or `["strip", "auto-orient", {"implode", "1"}]`.
## Email
### Pleroma.Emails.Mailer
* `adapter`: one of the mail adapters listed in [Swoosh readme](https://github.com/swoosh/swoosh#adapters), or `Swoosh.Adapters.Local` for in-memory mailbox.
* `api_key` / `password` and / or other adapter-specific settings, per the above documentation.
* `enabled`: Allows enable/disable send emails. Default: `false`.
An example for Sendgrid adapter:
```elixir
config :pleroma, Pleroma.Emails.Mailer,
enabled: true,
adapter: Swoosh.Adapters.Sendgrid,
api_key: "YOUR_API_KEY"
```
An example for SMTP adapter:
```elixir
config :pleroma, Pleroma.Emails.Mailer,
enabled: true,
adapter: Swoosh.Adapters.Mua,
relay: "smtp.gmail.com",
auth: [username: "YOUR_USERNAME@gmail.com", password: "YOUR_SMTP_PASSWORD"],
port: 465,
protocol: :ssl
```
An example for Mua adapter:
```elixir
config :pleroma, Pleroma.Emails.Mailer,
enabled: true,
adapter: Swoosh.Adapters.Mua,
relay: "mail.example.com",
port: 465,
auth: [
username: "YOUR_USERNAME@domain.tld",
password: "YOUR_SMTP_PASSWORD"
],
protocol: :ssl
```
### :email_notifications
Email notifications settings.
- digest - emails of "what you've missed" for users who have been
inactive for a while.
- active: globally enable or disable digest emails
- schedule: When to send digest email, in [crontab format](https://en.wikipedia.org/wiki/Cron).
"0 0 * * 0" is the default, meaning "once a week at midnight on Sunday morning"
- interval: Minimum interval between digest emails to one user
- inactivity_threshold: Minimum user inactivity threshold
### Pleroma.Emails.UserEmail
- `:logo` - a path to a custom logo. Set it to `nil` to use the default Pleroma logo.
- `:styling` - a map with color settings for email templates.
### Pleroma.Emails.NewUsersDigestEmail
- `:enabled` - a boolean, enables new users admin digest email when `true`. Defaults to `false`.
## Background jobs
### Oban
[Oban](https://github.com/sorentwo/oban) asynchronous job processor configuration.
Configuration options described in [Oban readme](https://github.com/sorentwo/oban#usage):
* `repo` - app's Ecto repo (`Pleroma.Repo`)
* `log` - logs verbosity
* `queues` - job queues (see below)
* `crontab` - periodic jobs, see [`Oban.Cron`](#obancron)
Pleroma has the following queues:
* `activity_expiration` - Activity expiration
* `federator_outgoing` - Outgoing federation
* `federator_incoming` - Incoming federation
* `mailer` - Email sender, see [`Pleroma.Emails.Mailer`](#pleromaemailsmailer)
* `transmogrifier` - Transmogrifier
* `web_push` - Web push notifications
* `scheduled_activities` - Scheduled activities, see [`Pleroma.ScheduledActivity`](#pleromascheduledactivity)
#### Oban.Cron
Pleroma has these periodic job workers:
* `Pleroma.Workers.Cron.DigestEmailsWorker` - digest emails for users with new mentions and follows
* `Pleroma.Workers.Cron.NewUsersDigestWorker` - digest emails for admins with new registrations
```elixir
config :pleroma, Oban,
repo: Pleroma.Repo,
verbose: false,
prune: {:maxlen, 1500},
queues: [
federator_incoming: 50,
federator_outgoing: 50
],
crontab: [
{"0 0 * * 0", Pleroma.Workers.Cron.DigestEmailsWorker},
{"0 0 * * *", Pleroma.Workers.Cron.NewUsersDigestWorker}
]
```
This config contains two queues: `federator_incoming` and `federator_outgoing`. Both have the number of max concurrent jobs set to `50`.
#### Migrating `pleroma_job_queue` settings
`config :pleroma_job_queue, :queues` is replaced by `config :pleroma, Oban, :queues` and uses the same format (keys are queues' names, values are max concurrent jobs numbers).
### :workers
Includes custom worker options not interpretable directly by `Oban`.
* `retries` — keyword lists where keys are `Oban` queues (see above) and values are numbers of max attempts for failed jobs.
Example:
```elixir
config :pleroma, :workers,
retries: [
federator_incoming: 5,
federator_outgoing: 5
]
```
#### Migrating `Pleroma.Web.Federator.RetryQueue` settings
* `max_retries` is replaced with `config :pleroma, :workers, retries: [federator_outgoing: 5]`
* `enabled: false` corresponds to `config :pleroma, :workers, retries: [federator_outgoing: 1]`
* deprecated options: `max_jobs`, `initial_timeout`
## :web_push_encryption, :vapid_details
Web Push Notifications configuration. You can use the mix task `mix web_push.gen.keypair` to generate it.
* ``subject``: a mailto link for the administrative contact. It’s best if this email is not a personal email address, but rather a group email so that if a person leaves an organization, is unavailable for an extended period, or otherwise can’t respond, someone else on the list can.
* ``public_key``: VAPID public key
* ``private_key``: VAPID private key
## :logger
* Logging to console/stdout is done by default, use `{ExSyslogger, :ex_syslogger}` to log to syslog
An example to enable ONLY ExSyslogger (f/ex in ``prod.secret.exs``) with info and debug suppressed:
```elixir
config :pleroma, :logger,
backends: [{ExSyslogger, :ex_syslogger}]
config :logger, default_handler: false
config :logger, :ex_syslogger,
level: :warning
```
Another example, keeping console output and adding the pid to syslog output:
```elixir
config :pleroma, :logger,
backends: [{ExSyslogger, :ex_syslogger}]
config :logger, :ex_syslogger,
level: :warning,
option: [:pid, :ndelay]
```
See: [logger’s documentation](https://hexdocs.pm/logger/Logger.html) and [ex_syslogger’s documentation](https://hexdocs.pm/ex_syslogger/)
An example of logging info to local syslog, but debug to console:
```elixir
config :pleroma, :logger,
backends: [{ExSyslogger, :ex_syslogger}]
config :logger, :ex_syslogger,
level: :info,
ident: "pleroma",
format: "$metadata[$level] $message"
config :logger, :default_handler,
level: :debug
config :logger, :default_formatter,
format: "\n$time $metadata[$level] $message\n",
metadata: [:request_id]
```
## Database options
### RUM indexing for full text search
* `rum_enabled`: If RUM indexes should be used. Defaults to `false`.
RUM indexes are an alternative indexing scheme that is not included in PostgreSQL by default. While they may eventually be mainlined, for now they have to be installed as a PostgreSQL extension from [https://github.com/postgrespro/rum](https://github.com/postgrespro/rum).
Their advantage over the standard GIN indexes is that they allow efficient ordering of search results by timestamp, which makes search queries a lot faster on larger servers, by one or two orders of magnitude. They take up around 3-4 times as much space as GIN indexes.
To enable them, both the `rum_enabled` flag has to be set and the following special migration has to be run:
* Source install:
- Stop Pleroma
- `mix ecto.migrate --migrations-path priv/repo/optional_migrations/rum_indexing/`
* OTP install:
- Stop Pleroma
- `pleroma_ctl migrate --migrations-path priv/repo/optional_migrations/rum_indexing/`
This will probably take a long time.
!!! note
It is recommended to `VACUUM FULL` the objects table after the migration has completed, to do that run:
```
# sudo -Hu postgres vacuumdb --full --analyze -t objects <pleroma DB name>
```
Now you can start Pleroma back up.
## Alternative client protocols
### BBS / SSH access
This feature has been removed from Pleroma core.
However, a client has been made and is available at https://git.pleroma.social/Duponin/sshocial.
### :gopher
* `enabled`: Enables the gopher interface
* `ip`: IP address to bind to
* `port`: Port to bind to
* `dstport`: Port advertised in urls (optional, defaults to `port`)
## Authentication
### :admin_token
Allows to set a token that can be used to authenticate with the admin api without using an actual user by giving it as the `admin_token` parameter or `x-admin-token` HTTP header. Example:
```elixir
config :pleroma, :admin_token, "somerandomtoken"
```
You can then do
```shell
curl "http://localhost:4000/api/v1/pleroma/admin/users/invites?admin_token=somerandomtoken"
```
or
```shell
curl -H "X-Admin-Token: somerandomtoken" "http://localhost:4000/api/v1/pleroma/admin/users/invites"
```
Warning: it's discouraged to use this feature because of the associated security risk: static / rarely changed instance-wide token is much weaker compared to email-password pair of a real admin user; consider using HTTP Basic Auth or OAuth-based authentication instead.
### :auth
Authentication / authorization settings.
* `auth_template`: authentication form template. By default it's `show.html` which corresponds to `lib/pleroma/web/templates/o_auth/o_auth/show.html.eex`.
* `oauth_consumer_template`: OAuth consumer mode authentication form template. By default it's `consumer.html` which corresponds to `lib/pleroma/web/templates/o_auth/o_auth/consumer.html.eex`.
* `oauth_consumer_strategies`: the list of enabled OAuth consumer strategies; by default it's set by `OAUTH_CONSUMER_STRATEGIES` environment variable. Each entry in this space-delimited string should be of format `<strategy>` or `<strategy>:<dependency>` (e.g. `twitter` or `keycloak:ueberauth_keycloak_strategy` in case dependency is named differently than `ueberauth_<strategy>`).
### Pleroma.Web.Auth.Authenticator
* `Pleroma.Web.Auth.PleromaAuthenticator`: default database authenticator.
* `Pleroma.Web.Auth.LDAPAuthenticator`: LDAP authentication.
### :ldap
Use LDAP for user authentication. When a user logs in to the Pleroma
instance, the name and password will be verified by trying to authenticate
(bind) to an LDAP server. If a user exists in the LDAP directory but there
is no account with the same name yet on the Pleroma instance then a new
Pleroma account will be created with the same name as the LDAP user name.
* `enabled`: enables LDAP authentication
* `host`: LDAP server hostname
* `port`: LDAP port, e.g. 389 or 636
* `ssl`: true to use implicit SSL/TLS, usually port 636
* `sslopts`: additional SSL options
* `tls`: true to use explicit TLS (STARTTLS), usually port 389
* `tlsopts`: additional TLS options
* `base`: LDAP base, e.g. "dc=example,dc=com"
* `uid`: LDAP attribute name to authenticate the user, e.g. when "cn", the filter will be "cn=username,base"
* `cacertfile`: Path to alternate CA root certificates file
Note, if your LDAP server is an Active Directory server the correct value is commonly `uid: "cn"`, but if you use an
OpenLDAP server the value may be `uid: "uid"`.
### :oauth2 (Pleroma as OAuth 2.0 provider settings)
OAuth 2.0 provider settings:
* `token_expires_in` - The lifetime in seconds of the access token.
* `issue_new_refresh_token` - Keeps old refresh token or generate new refresh token when to obtain an access token.
* `clean_expired_tokens` - Enable a background job to clean expired oauth tokens. Defaults to `false`.
OAuth 2.0 provider and related endpoints:
* `POST /api/v1/apps` creates client app basing on provided params.
* `GET/POST /oauth/authorize` renders/submits authorization form.
* `POST /oauth/token` creates/renews OAuth token.
* `POST /oauth/revoke` revokes provided OAuth token.
* `GET /api/v1/accounts/verify_credentials` (with proper `Authorization` header or `access_token` URI param) returns user info on requester (with `acct` field containing local nickname and `fqn` field containing fully-qualified nickname which could generally be used as email stub for OAuth software that demands email field in identity endpoint response, like Peertube).
### OAuth consumer mode
OAuth consumer mode allows sign in / sign up via external OAuth providers (e.g. Twitter, Facebook, Google, Microsoft, etc.).
Implementation is based on Ueberauth; see the list of [available strategies](https://github.com/ueberauth/ueberauth/wiki/List-of-Strategies).
!!! note
Each strategy is shipped as a separate dependency; in order to get the strategies, run `OAUTH_CONSUMER_STRATEGIES="..." mix deps.get`, e.g. `OAUTH_CONSUMER_STRATEGIES="twitter facebook google microsoft" mix deps.get`. The server should also be started with `OAUTH_CONSUMER_STRATEGIES="..." mix phx.server` in case you enable any strategies.
!!! note
Each strategy requires separate setup (on external provider side and Pleroma side). Below are the guidelines on setting up most popular strategies.
!!! note
Make sure that `"SameSite=Lax"` is set in `extra_cookie_attrs` when you have this feature enabled. OAuth consumer mode will not work with `"SameSite=Strict"`
* For Twitter, [register an app](https://developer.twitter.com/en/apps), configure callback URL to https://<your_host>/oauth/twitter/callback
* For Facebook, [register an app](https://developers.facebook.com/apps), configure callback URL to https://<your_host>/oauth/facebook/callback, enable Facebook Login service at https://developers.facebook.com/apps/<app_id>/fb-login/settings/
* For Google, [register an app](https://console.developers.google.com), configure callback URL to https://<your_host>/oauth/google/callback
* For Microsoft, [register an app](https://portal.azure.com), configure callback URL to https://<your_host>/oauth/microsoft/callback
Once the app is configured on external OAuth provider side, add app's credentials and strategy-specific settings (if any — e.g. see Microsoft below) to `config/prod.secret.exs`,
per strategy's documentation (e.g. [ueberauth_twitter](https://github.com/ueberauth/ueberauth_twitter)). Example config basing on environment variables:
```elixir
# Twitter
config :ueberauth, Ueberauth.Strategy.Twitter.OAuth,
consumer_key: System.get_env("TWITTER_CONSUMER_KEY"),
consumer_secret: System.get_env("TWITTER_CONSUMER_SECRET")
# Facebook
config :ueberauth, Ueberauth.Strategy.Facebook.OAuth,
client_id: System.get_env("FACEBOOK_APP_ID"),
client_secret: System.get_env("FACEBOOK_APP_SECRET"),
redirect_uri: System.get_env("FACEBOOK_REDIRECT_URI")
# Google
config :ueberauth, Ueberauth.Strategy.Google.OAuth,
client_id: System.get_env("GOOGLE_CLIENT_ID"),
client_secret: System.get_env("GOOGLE_CLIENT_SECRET"),
redirect_uri: System.get_env("GOOGLE_REDIRECT_URI")
# Microsoft
config :ueberauth, Ueberauth.Strategy.Microsoft.OAuth,
client_id: System.get_env("MICROSOFT_CLIENT_ID"),
client_secret: System.get_env("MICROSOFT_CLIENT_SECRET")
config :ueberauth, Ueberauth,
providers: [
microsoft: {Ueberauth.Strategy.Microsoft, [callback_params: []]}
]
# Keycloak
# Note: make sure to add `keycloak:ueberauth_keycloak_strategy` entry to `OAUTH_CONSUMER_STRATEGIES` environment variable
keycloak_url = "https://publicly-reachable-keycloak-instance.org:8080"
config :ueberauth, Ueberauth.Strategy.Keycloak.OAuth,
client_id: System.get_env("KEYCLOAK_CLIENT_ID"),
client_secret: System.get_env("KEYCLOAK_CLIENT_SECRET"),
site: keycloak_url,
authorize_url: "#{keycloak_url}/auth/realms/master/protocol/openid-connect/auth",
token_url: "#{keycloak_url}/auth/realms/master/protocol/openid-connect/token",
userinfo_url: "#{keycloak_url}/auth/realms/master/protocol/openid-connect/userinfo",
token_method: :post
config :ueberauth, Ueberauth,
providers: [
keycloak: {Ueberauth.Strategy.Keycloak, [uid_field: :email]}
]
```
## Link parsing
### :uri_schemes
* `valid_schemes`: List of the scheme part that is considered valid to be an URL.
### Pleroma.Formatter
Configuration for Pleroma's link formatter which parses mentions, hashtags, and URLs.
* `class` - specify the class to be added to the generated link (default: `false`)
* `rel` - specify the rel attribute (default: `ugc`)
* `new_window` - adds `target="_blank"` attribute (default: `false`)
* `truncate` - Set to a number to truncate URLs longer then the number. Truncated URLs will end in `...` (default: `false`)
* `strip_prefix` - Strip the scheme prefix (default: `false`)
* `extra` - link URLs with rarely used schemes (magnet, ipfs, irc, etc.) (default: `true`)
* `validate_tld` - Set to false to disable TLD validation for URLs/emails. Can be set to :no_scheme to validate TLDs only for urls without a scheme (e.g `example.com` will be validated, but `http://example.loki` won't) (default: `:no_scheme`)
Example:
```elixir
config :pleroma, Pleroma.Formatter,
class: false,
rel: "ugc",
new_window: false,
truncate: false,
strip_prefix: false,
extra: true,
validate_tld: :no_scheme
```
## Custom Runtime Modules (`:modules`)
* `runtime_dir`: A path to custom Elixir modules (such as MRF policies).
## :configurable_from_database
Boolean, enables/disables in-database configuration. Read [Transferring the config to/from the database](../administration/CLI_tasks/config.md) for more information.
## :database_config_whitelist
List of valid configuration sections which are allowed to be configured from the
database. Settings stored in the database before the whitelist is configured are
still applied. Consider running the `mix pleroma.config filter_whitelisted` task
after updating the whitelist. Read [Remove non-whitelisted configs from the database](../administration/CLI_tasks/config.md#remove-non-whitelisted-configs-from-the-database)
for more information.
Example:
```elixir
config :pleroma, :database_config_whitelist, [
{:pleroma, :instance},
{:pleroma, Pleroma.Web.Metadata},
{:auto_linker}
]
```
### Multi-factor authentication - :two_factor_authentication
* `totp` - a list containing TOTP configuration
- `digits` - Determines the length of a one-time pass-code in characters. Defaults to 6 characters.
- `period` - a period for which the TOTP code will be valid in seconds. Defaults to 30 seconds.
* `backup_codes` - a list containing backup codes configuration
- `number` - number of backup codes to generate.
- `length` - backup code length. Defaults to 16 characters.
## Restrict entities access for unauthenticated users
### :restrict_unauthenticated
Restrict access for unauthenticated users to timelines (public and federated), user profiles and statuses.
* `timelines`: public and federated timelines
* `local`: public timeline
* `federated`: federated timeline (includes public timeline)
* `profiles`: user profiles
* `local`
* `remote`
* `activities`: statuses
* `local`
* `remote`
Note: when `:instance, :public` is set to `false`, all `:restrict_unauthenticated` items be effectively set to `true` by default. If you'd like to allow unauthenticated access to specific API endpoints on a private instance, please explicitly set `:restrict_unauthenticated` to non-default value in `config/prod.secret.exs`.
Note: setting `restrict_unauthenticated/timelines/local` to `true` has no practical sense if `restrict_unauthenticated/timelines/federated` is set to `false` (since local public activities will still be delivered to unauthenticated users as part of federated timeline).
## Pleroma.Web.ApiSpec.CastAndValidate
* `:strict` a boolean, enables strict input validation (useful in development, not recommended in production). Defaults to `false`.
## :instances_favicons
Control favicons for instances.
* `enabled`: Allow/disallow displaying and getting instances favicons
## :retention
Bounded retention of remote content. Local posts are your data; remote posts are a cache of somebody else's, and this treats them that way. When enabled, a cron worker (`Pleroma.Workers.Cron.RetentionWorker`, hourly by default) evicts a batch of remote threads that nobody local has interacted with and that have been quiet for longer than `remote_post_retention_days` (see [`:instance`](#instance)). The database then plateaus instead of growing without limit.
-A thread is never evicted if it contains a local post, reply, favourite, repeat or reaction, a post bookmarked by a local user, a post addressed to a local user (direct messages and mentions), a post a local user still has a notification for, or a post that was reported. Relationship activities (follows, blocks) and chat messages are never touched. Evicted posts are refetched on demand when somebody opens them again, as with `prune_objects`.
-
-!!! warning
- A local quote post does not pin the post it quotes, since quotes do not share the quoted thread's context. Old quoted remote posts can be evicted and will be refetched when the quote is viewed.
+A thread is never evicted if it contains a local post, reply, favourite, repeat or reaction, a post bookmarked by a local user, a post addressed to a local user (direct messages and mentions), a post quoted by a local post, a post a local user still has a notification for, or a post that was reported. Relationship activities (follows, blocks) and chat messages are never touched. Evicted posts are refetched on demand when somebody opens them again, as with `prune_objects`.
* `enabled`: Run the worker. Defaults to `false`.
* `max_objects`: Optional watermark. When the `objects` table is estimated to hold more rows than this, the oldest unpinned remote threads that have been quiet for more than a day are evicted, whatever `remote_post_retention_days` says. Defaults to `nil` (age only).
* `batch_size`: Number of activities walked per run. Their threads are checked and, if unpinned, evicted. Defaults to `50000`.
* `keep_non_public`: Also keep threads that contain a non-public post. Defaults to `false`.
!!! note
The worker walks the `activities` table in id order, `batch_size` activities per run, up to the id that corresponds to the retention deadline, and remembers its position in the `retention_cursors` table. Each activity is visited once, when it becomes old enough, and only the threads it belongs to are checked through the context index. Over the `max_objects` watermark a second walk with its own cursor runs up to one day ago instead. The first run starts from the oldest activity; on an instance that never pruned this backlog takes many runs, so run `prune_objects --keep-threads --prune-orphaned-activities` once for the initial cut. Threads that were kept and later lost their pin (removed bookmarks, cleared notifications of other kinds) are only revisited on a fresh walk: delete the row from `retention_cursors` to restart from the beginning. Space freed by eviction is reused by new rows; it is not handed back to the filesystem unless you `VACUUM FULL`.
## Pleroma.User.Backup
!!! note
Requires enabled email
* `:purge_after_days` an integer, remove backup achieves after N days.
* `:limit_days` an integer, limit user to export not more often than once per N days.
* `:dir` a string with a path to backup temporary directory or `nil` to let Pleroma choose temporary directory in the following order:
1. the directory named by the TMPDIR environment variable
2. the directory named by the TEMP environment variable
3. the directory named by the TMP environment variable
4. C:\TMP on Windows or /tmp on Unix-like operating systems
5. as a last resort, the current working directory
* `:timeout` an integer representing seconds
## Frontend management
Frontends in Pleroma are swappable - you can specify which one to use here.
You can set a frontends for the key `primary` and `admin` and the options of `name` and `ref`. This will then make Pleroma serve the frontend from a folder constructed by concatenating the instance static path, `frontends` and the name and ref.
The key `primary` refers to the frontend that will be served by default for general requests. The key `admin` refers to the frontend that will be served at the `/pleroma/admin` path.
If you don't set anything here, the bundled frontends will be used.
Example:
```
config :pleroma, :frontends,
primary: %{
"name" => "pleroma",
"ref" => "stable"
},
admin: %{
"name" => "admin",
"ref" => "develop"
}
```
This would serve the frontend from the the folder at `$instance_static/frontends/pleroma/stable`. You have to copy the frontend into this folder yourself. You can choose the name and ref any way you like, but they will be used by mix tasks to automate installation in the future, the name referring to the project and the ref referring to a commit.
## Ephemeral activities (Pleroma.Workers.PurgeExpiredActivity)
Settings to enable and configure expiration for ephemeral activities
* `:enabled` - enables ephemeral activities creation
* `:min_lifetime` - minimum lifetime for ephemeral activities (in seconds). Default: 10 minutes.
## ConcurrentLimiter
Settings to restrict concurrently running jobs. Jobs which can be configured:
* `Pleroma.Web.RichMedia.Helpers` - generating link previews of URLs in activities
* `Pleroma.Web.ActivityPub.MRF.MediaProxyWarmingPolicy` - warming remote media cache via MediaProxyWarmingPolicy
Each job has these settings:
* `:max_running` - max concurrently runnings jobs
* `:max_waiting` - max waiting jobs
diff --git a/lib/pleroma/retention.ex b/lib/pleroma/retention.ex
index 9fe15d0b8..563667bb1 100644
--- a/lib/pleroma/retention.ex
+++ b/lib/pleroma/retention.ex
@@ -1,460 +1,473 @@
# Pleroma: A lightweight social networking server
# Copyright © 2017-2022 Pleroma Authors <https://pleroma.social/>
# SPDX-License-Identifier: AGPL-3.0-only
defmodule Pleroma.Retention do
@moduledoc """
Bounded retention of remote content.
Local content is data; remote content is a cache of somebody else's data.
This module treats it that way: a remote thread is evictable when nobody
local has interacted with it and it has been quiet for longer than
`remote_post_retention_days`, or for a day when the objects table has grown
past the configured watermark. Local posts are never evicted.
A thread is pinned (never evicted) if any activity in it
* is local (posts, replies, favourites, repeats and reactions by local
users all carry the thread's context),
* is bookmarked by a local user,
* is addressed to a local user (direct messages and mentions),
+ * is a post quoted by a local post,
* still has a notification for a local user,
* is a report (Flag),
or if any object in it is referenced by a report.
Only activities that carry a context take part; relationship activities
such as Follow or Block and chat messages have none and are never touched.
## How a run works
Grouping the whole activities table by thread is far too expensive to do
every hour on a large instance, so `run/1` walks the table incrementally
instead. Activity ids are ordered by creation time, so a run scans the next
`batch_size` activities after a persisted cursor (`Pleroma.Retention.Cursor`)
up to the id that corresponds to the retention deadline, collects the
threads they belong to, and verifies just those threads through the
`(type, context)` index. Threads that pass are evicted; the cursor moves on.
Every activity is thus visited exactly once, when it becomes old enough,
and a thread is re-checked whenever another of its activities ages out.
A thread that was kept because an activity's `updated_at` was bumped
after insertion is only re-checked on the next full walk, i.e. after a
cursor reset. Over the object watermark a second walk, with its own
cursor, runs up to a day ago instead of the deadline.
The first run starts from the oldest activity, which doubles as the
initial cut on an instance that never pruned. `prune_objects --keep-threads`
does the same job in one go and is faster for that purpose.
"""
import Ecto.Query
alias Pleroma.Activity
alias Pleroma.Bookmark
alias Pleroma.Config
alias Pleroma.Notification
alias Pleroma.Object
alias Pleroma.Repo
alias Pleroma.Retention.Cursor
require Logger
require Pleroma.Constants
@type stats :: %{
contexts: non_neg_integer(),
objects: non_neg_integer(),
activities: non_neg_integer()
}
@empty_stats %{contexts: 0, objects: 0, activities: 0}
@default_batch_size 50_000
# Over the watermark a thread must still have been quiet this long, so
# posts people are reading right now are not evicted from under them.
@watermark_min_age 86_400
# Activities per verification chunk; the cursor is persisted after each.
@chunk 5_000
# Activity types that may carry a thread's context. Restricting the
# per-thread lookups to these lets them use the (type, context) index.
#
# Invariant: every type that *adopts* another activity's context must be
# listed, or its thread's local pins become invisible to the walk and the
# thread gets evicted. Today only Create, Announce, Like and EmojiReact do
# (see Pleroma.Web.ActivityPub.Builder), and EmojiReaction, the name
# EmojiReact had in older data. The others carry a context of their own (or
# did in older data) and are listed so they get evicted rather than left
# behind. The regression test "keeps threads pinned by each local
# interaction type" guards this.
@context_types ~w(Create Announce Like EmojiReact EmojiReaction Update Delete Undo Flag Listen Move)
@doc """
Runs one eviction batch using the `:retention` config and returns what was removed.
"""
@spec run(keyword()) :: stats()
def run(config \\ Config.get(:retention, [])) do
budget = config[:batch_size] || @default_batch_size
# Over the watermark every unpinned thread older than a day is fair game,
# so a separate walk goes up to that point. It must not share the cursor
# with the age-based walk, or that one would be left behind the deadline.
{cursor, deadline} =
if over_watermark?(config[:max_objects]),
do: {"activities_watermark", watermark_deadline()},
else: {"activities", deadline()}
boundary = id_at(deadline)
opts = [
keep_non_public: config[:keep_non_public],
skip_reported: reported_contexts()
]
from = Cursor.get(cursor)
rows = candidates(from, boundary, budget)
stats =
rows
|> Enum.chunk_every(@chunk)
|> Enum.reduce(@empty_stats, fn chunk, acc ->
stats = chunk |> contexts_of() |> evict_threads(deadline, opts)
# Persist progress per chunk, so a killed or crashed run resumes
# where it stopped instead of re-walking the whole range.
Cursor.put(cursor, chunk |> List.last() |> elem(0))
Map.merge(acc, stats, fn _key, x, y -> x + y end)
end)
# Fewer rows than asked for: everything up to the boundary is done.
if length(rows) < budget and after?(boundary, from), do: Cursor.put(cursor, boundary)
stats
end
# The next `budget` activities after `from`, up to `boundary`.
defp candidates(from, boundary, budget) do
Activity
|> where([a], a.id <= ^boundary)
|> maybe_after(from)
|> order_by([a], asc: a.id)
|> limit(^budget)
|> select([a], {a.id, a.local, fragment("? ->> 'context'", a.data)})
|> Repo.all(timeout: :infinity)
end
defp maybe_after(query, nil), do: query
defp maybe_after(query, from), do: where(query, [a], a.id > ^from)
defp contexts_of(rows) do
for {_id, false, context} when is_binary(context) <- rows, uniq: true, do: context
end
defp after?(_id, nil), do: true
defp after?(id, other), do: FlakeId.from_string(id) > FlakeId.from_string(other)
# Verifies the given threads against the pin rules and evicts those that
# pass. The transaction makes a chunk all-or-nothing but does not lock the
# threads: a local interaction landing between check and delete survives,
# pins the thread from then on, and its object is refetched when needed.
defp evict_threads([], _deadline, _opts), do: @empty_stats
defp evict_threads(contexts, deadline, opts) do
query =
deadline
|> evictable_contexts_query(Keyword.put(opts, :contexts, contexts))
|> select([a], %{
context: fragment("? ->> 'context'", a.data),
activity_ids: type(fragment("array_agg(?)", a.id), {:array, FlakeId.Ecto.CompatType}),
object_ids: fragment("array_agg(associated_object_id(?))", a.data)
})
{:ok, stats} =
Repo.transaction(
fn ->
query
|> Repo.all(timeout: :infinity)
|> evict()
end,
timeout: :infinity
)
stats
end
# The activity id a flake would have received at the given time. Ids at or
# below it belong to activities created before then.
defp id_at(%NaiveDateTime{} = time) do
ms = time |> DateTime.from_naive!("Etc/UTC") |> DateTime.to_unix(:millisecond) |> max(0)
FlakeId.to_string(<<ms::integer-size(64), 0::integer-size(64)>>)
end
@doc """
The point in time before which a quiet remote thread becomes evictable.
"""
@spec deadline() :: NaiveDateTime.t()
def deadline do
days = Config.get([:instance, :remote_post_retention_days])
NaiveDateTime.add(NaiveDateTime.utc_now(), -days * 86_400)
end
# The watermark deadline is never further back than the normal one.
defp watermark_deadline do
min_age = NaiveDateTime.add(NaiveDateTime.utc_now(), -@watermark_min_age)
Enum.max([deadline(), min_age], NaiveDateTime)
end
@doc """
Activities grouped by thread context, restricted to threads that may be evicted.
The query has no `select`; callers pick what they need from the group (the
context alone for `IN (subquery)`, or aggregates of ids for deletion).
Options:
* `:keep_non_public` - also pin threads containing a non-public post
* `:contexts` - only look at these threads, using the `(type, context)`
index; without it the whole table is grouped
* `:skip_reported` - a `MapSet` of reported contexts already removed from
`:contexts` by the caller, so the subquery for them can be skipped
"""
@spec evictable_contexts_query(NaiveDateTime.t(), keyword()) :: Ecto.Query.t()
def evictable_contexts_query(deadline, opts \\ []) do
contexts =
case {Keyword.get(opts, :contexts), Keyword.get(opts, :skip_reported)} do
{nil, _} -> nil
{contexts, nil} -> contexts
{contexts, reported} -> Enum.reject(contexts, &MapSet.member?(reported, &1))
end
Activity
|> join(:left, [a], b in Bookmark, on: a.id == b.activity_id)
|> join(:left, [a], n in Notification, on: a.id == n.activity_id)
|> where([a], not is_nil(fragment("? ->> 'context'", a.data)))
|> maybe_exclude_reported(is_nil(Keyword.get(opts, :skip_reported)))
|> maybe_only_contexts(contexts)
|> group_by([a], fragment("? ->> 'context'", a.data))
|> having([a], max(a.updated_at) < ^deadline)
|> having([a], not fragment("bool_or(?)", a.local))
|> having([_, b], fragment("max(?::text) is null", b.id))
|> having([_, _, n], fragment("max(?) is null", n.id))
|> having(
[a],
not fragment(
"bool_or(EXISTS (SELECT 1 FROM unnest(?) AS r WHERE left(r, ?) = ?))",
a.recipients,
^local_prefix_length(),
^local_prefix()
)
)
+ # A local quote has a context of its own, so look up quotes of the posts
+ # here; the objects_quote_url index is on the jsonb value.
+ |> having(
+ [a],
+ not fragment(
+ "bool_or(? ->> 'type' = 'Create' AND EXISTS (SELECT 1 FROM objects q WHERE q.data -> 'quoteUrl' = to_jsonb(associated_object_id(?)) AND left(q.data ->> 'actor', ?) = ?))",
+ a.data,
+ a.data,
+ ^local_prefix_length(),
+ ^local_prefix()
+ )
+ )
|> maybe_keep_non_public(Keyword.get(opts, :keep_non_public, false) == true)
end
defp maybe_exclude_reported(query, false), do: query
defp maybe_exclude_reported(query, true) do
where(
query,
[a],
fragment("? ->> 'context'", a.data) not in subquery(reported_contexts_query())
)
end
@doc "Contexts of reported objects and of the reports themselves."
@spec reported_contexts() :: MapSet.t(String.t())
def reported_contexts do
reported_contexts_query() |> Repo.all(timeout: :infinity) |> MapSet.new()
end
defp maybe_only_contexts(query, nil), do: query
defp maybe_only_contexts(query, contexts) do
query
|> where([a], fragment("? ->> 'type'", a.data) in ^@context_types)
|> where([a], fragment("? ->> 'context'", a.data) in ^contexts)
end
defp maybe_keep_non_public(query, false), do: query
defp maybe_keep_non_public(query, true) do
having(
query,
[a],
not fragment(
# Posts (checked on Create Activity) is non-public
"bool_or((not(?->'to' \\? ? OR ?->'cc' \\? ?)) and ? ->> 'type' = 'Create')",
a.data,
^Pleroma.Constants.as_public(),
a.data,
^Pleroma.Constants.as_public(),
a.data
)
)
end
# Contexts touched by reports: the contexts of the reported objects, plus the
# Flag activities' own contexts (a report gets a fresh context of its own).
# Flag objects are a list mixing actor ids, object ids and embedded copies
# of objects with an "id"; a single object instead of a list is tolerated.
defp reported_contexts_query do
flags = where(Activity, [f], fragment("? ->> 'type' = 'Flag'", f.data))
reported_objects =
flags
|> join(
:inner,
[f],
fo in fragment(
"jsonb_array_elements(CASE WHEN jsonb_typeof(? -> 'object') = 'array' THEN ? -> 'object' ELSE jsonb_build_array(? -> 'object') END)",
f.data,
f.data,
f.data
),
on: true
)
|> join(:inner, [_, fo], o in Object,
on: fragment("? ->> 'id' = coalesce(? ->> 'id', ? #>> '{}')", o.data, fo.value, fo.value)
)
|> where([_, _, o], not is_nil(fragment("? ->> 'context'", o.data)))
|> select([_, _, o], fragment("? ->> 'context'", o.data))
flags
|> where([f], not is_nil(fragment("? ->> 'context'", f.data)))
|> select([f], fragment("? ->> 'context'", f.data))
|> union(^reported_objects)
end
@doc """
Deletes the given threads. Each entry carries the context, activity ids and
object ids collected per thread, so deletion uses primary key and ap_id
indexes only.
An object is only deleted if it belongs to one of the threads itself: a
Like or Announce may carry a context other than its object's, and that
object's own thread can be pinned.
"""
@spec evict([
%{context: String.t(), activity_ids: [String.t()], object_ids: [String.t() | nil]}
]) :: stats()
def evict([]), do: @empty_stats
def evict(threads) do
activity_ids = Enum.flat_map(threads, & &1.activity_ids)
object_ids = threads |> Enum.flat_map(& &1.object_ids) |> Enum.reject(&is_nil/1)
contexts = Enum.map(threads, & &1.context)
{:ok, stats} =
Repo.transaction(
fn ->
# hashtags_objects rows go with the objects, so collect them first
hashtag_ids =
from(ho in "hashtags_objects",
join: o in Object,
on: o.id == ho.object_id,
where: fragment("? ->> 'id'", o.data) in ^object_ids,
where: fragment("? ->> 'context'", o.data) in ^contexts,
distinct: true,
select: ho.hashtag_id
)
|> Repo.all(timeout: :infinity)
{objects, deleted_ap_ids} =
Object
|> where([o], fragment("? ->> 'id'", o.data) in ^object_ids)
|> where([o], fragment("? ->> 'context'", o.data) in ^contexts)
|> where(
[o],
fragment(
"left(? ->> 'actor', ?) != ?",
o.data,
^local_prefix_length(),
^local_prefix()
)
)
|> select([o], fragment("? ->> 'id'", o.data))
|> Repo.delete_all(timeout: :infinity)
{activities, _} =
Activity
|> where([a], a.id in ^activity_ids)
|> where([a], a.local == false)
|> where([a], fragment("? ->> 'type' != 'Flag'", a.data))
|> Repo.delete_all(timeout: :infinity)
# Remote activities outside the thread that point at objects we just
# removed (e.g. Deletes carry no context). Flags are kept as they
# may have report notes attached.
{referencing, _} =
Activity
|> where([a], a.local == false)
|> where([a], fragment("associated_object_id(?)", a.data) in ^deleted_ap_ids)
|> where([a], fragment("? ->> 'type' != 'Flag'", a.data))
|> Repo.delete_all(timeout: :infinity)
delete_unused_hashtags(hashtag_ids)
%{
contexts: length(threads),
objects: objects,
activities: activities + referencing
}
end,
timeout: :infinity
)
stats
end
@doc """
Whether the objects table is estimated to hold more rows than `max_objects`.
"""
@spec over_watermark?(non_neg_integer() | nil) :: boolean()
def over_watermark?(nil), do: false
def over_watermark?(max_objects), do: estimated_object_count() > max_objects
@doc """
Row count of the objects table from planner statistics, falling back to an
exact count when the table has never been analyzed.
"""
@spec estimated_object_count() :: non_neg_integer()
def estimated_object_count do
%{rows: [[estimate]]} =
Repo.query!("SELECT reltuples::bigint FROM pg_class WHERE oid = 'objects'::regclass")
if estimate > 0, do: estimate, else: Repo.aggregate(Object, :count, timeout: :infinity)
end
@unused_hashtags """
DELETE FROM hashtags AS ht
WHERE NOT EXISTS (
SELECT 1 FROM hashtags_objects hto
WHERE ht.id = hto.hashtag_id
)
AND NOT EXISTS (
SELECT 1 FROM user_follows_hashtag ufh
WHERE ht.id = ufh.hashtag_id
)
"""
@doc """
Removes hashtags no longer attached to any object and not followed by anyone,
either all of them or only those among the given ids.
"""
@spec delete_unused_hashtags([integer()] | nil) :: :ok
def delete_unused_hashtags(ids \\ nil)
def delete_unused_hashtags([]), do: :ok
def delete_unused_hashtags(nil) do
Repo.query!(@unused_hashtags, [], timeout: :infinity)
:ok
end
def delete_unused_hashtags(ids) do
Repo.query!(@unused_hashtags <> "AND ht.id = ANY($1)", [ids], timeout: :infinity)
:ok
end
# Actor ids of local users start with the instance url; an exact prefix
# comparison avoids LIKE wildcards and host/port ambiguity.
defp local_prefix, do: Pleroma.Web.Endpoint.url() <> "/"
defp local_prefix_length, do: String.length(local_prefix())
end
diff --git a/test/mix/tasks/pleroma/database_test.exs b/test/mix/tasks/pleroma/database_test.exs
index 8f48d02de..31b414cc8 100644
--- a/test/mix/tasks/pleroma/database_test.exs
+++ b/test/mix/tasks/pleroma/database_test.exs
@@ -1,725 +1,748 @@
# Pleroma: A lightweight social networking server
# Copyright © 2017-2022 Pleroma Authors <https://pleroma.social/>
# SPDX-License-Identifier: AGPL-3.0-only
defmodule Mix.Tasks.Pleroma.DatabaseTest do
use Pleroma.DataCase, async: false
use Oban.Testing, repo: Pleroma.Repo
alias Pleroma.Activity
alias Pleroma.Bookmark
alias Pleroma.Hashtag
alias Pleroma.Object
alias Pleroma.Repo
alias Pleroma.User
alias Pleroma.Web.CommonAPI
import Pleroma.Factory
setup_all do
Mix.shell(Mix.Shell.Process)
on_exit(fn ->
Mix.shell(Mix.Shell.IO)
end)
:ok
end
describe "running remove_embedded_objects" do
test "it replaces objects with references" do
user = insert(:user)
{:ok, activity} = CommonAPI.post(user, %{status: "test"})
new_data = Map.put(activity.data, "object", activity.object.data)
{:ok, activity} =
activity
|> Activity.change(%{data: new_data})
|> Repo.update()
assert is_map(activity.data["object"])
Mix.Tasks.Pleroma.Database.run(["remove_embedded_objects"])
activity = Activity.get_by_id_with_object(activity.id)
assert is_binary(activity.data["object"])
end
end
describe "prune_objects" do
setup do
deadline = Pleroma.Config.get([:instance, :remote_post_retention_days]) + 1
old_insert_date =
Timex.now()
|> Timex.shift(days: -deadline)
|> Timex.to_naive_datetime()
|> NaiveDateTime.truncate(:second)
%{old_insert_date: old_insert_date}
end
test "it prunes old objects from the database", %{old_insert_date: old_insert_date} do
insert(:note)
%{id: note_remote_public_id} =
:note
|> insert()
|> Ecto.Changeset.change(%{updated_at: old_insert_date})
|> Repo.update!()
note_remote_non_public =
%{id: note_remote_non_public_id, data: note_remote_non_public_data} =
:note
|> insert()
note_remote_non_public
|> Ecto.Changeset.change(%{
updated_at: old_insert_date,
data: note_remote_non_public_data |> update_in(["to"], fn _ -> [] end)
})
|> Repo.update!()
assert length(Repo.all(Object)) == 3
Mix.Tasks.Pleroma.Database.run(["prune_objects"])
assert length(Repo.all(Object)) == 1
refute Object.get_by_id(note_remote_public_id)
refute Object.get_by_id(note_remote_non_public_id)
end
test "it cleans up bookmarks", %{old_insert_date: old_insert_date} do
user = insert(:user)
{:ok, old_object_activity} = CommonAPI.post(user, %{status: "yadayada"})
Repo.one(Object)
|> Ecto.Changeset.change(%{updated_at: old_insert_date})
|> Repo.update!()
{:ok, new_object_activity} = CommonAPI.post(user, %{status: "yadayada"})
{:ok, _} = Bookmark.create(user.id, old_object_activity.id)
{:ok, _} = Bookmark.create(user.id, new_object_activity.id)
assert length(Repo.all(Object)) == 2
assert length(Repo.all(Bookmark)) == 2
Mix.Tasks.Pleroma.Database.run(["prune_objects"])
assert length(Repo.all(Object)) == 1
assert length(Repo.all(Bookmark)) == 1
refute Bookmark.get(user.id, old_object_activity.id)
end
test "with the --keep-non-public option it still keeps non-public posts even if they are not local",
%{old_insert_date: old_insert_date} do
insert(:note)
%{id: note_remote_id} =
:note
|> insert()
|> Ecto.Changeset.change(%{updated_at: old_insert_date})
|> Repo.update!()
note_remote_non_public =
%{data: note_remote_non_public_data} =
:note
|> insert()
note_remote_non_public
|> Ecto.Changeset.change(%{
updated_at: old_insert_date,
data: note_remote_non_public_data |> update_in(["to"], fn _ -> [] end)
})
|> Repo.update!()
assert length(Repo.all(Object)) == 3
Mix.Tasks.Pleroma.Database.run(["prune_objects", "--keep-non-public"])
assert length(Repo.all(Object)) == 2
refute Object.get_by_id(note_remote_id)
end
test "with the --keep-threads and --keep-non-public option it keeps old threads with non-public replies even if the interaction is not local",
%{old_insert_date: old_insert_date} do
# For non-public we only check Create Activities because only these are relevant for threads
# Flags are always non-public, Announces from relays can be non-public...
remote_user1 = insert(:user, local: false)
remote_user2 = insert(:user, local: false)
# Old remote non-public reply (should be kept)
{:ok, old_remote_post1_activity} =
CommonAPI.post(remote_user1, %{status: "some thing", local: false})
old_remote_post1_activity
|> Ecto.Changeset.change(%{local: false, updated_at: old_insert_date})
|> Repo.update!()
{:ok, old_remote_non_public_reply_activity} =
CommonAPI.post(remote_user2, %{
status: "some reply",
in_reply_to_status_id: old_remote_post1_activity.id
})
old_remote_non_public_reply_activity
|> Ecto.Changeset.change(%{
local: false,
updated_at: old_insert_date,
data: old_remote_non_public_reply_activity.data |> update_in(["to"], fn _ -> [] end)
})
|> Repo.update!()
# Old remote non-public Announce (should be removed)
{:ok, old_remote_post2_activity = %{data: %{"object" => old_remote_post2_id}}} =
CommonAPI.post(remote_user1, %{status: "some thing", local: false})
old_remote_post2_activity
|> Ecto.Changeset.change(%{local: false, updated_at: old_insert_date})
|> Repo.update!()
{:ok, old_remote_non_public_repeat_activity} =
CommonAPI.repeat(old_remote_post2_activity.id, remote_user2)
old_remote_non_public_repeat_activity
|> Ecto.Changeset.change(%{
local: false,
updated_at: old_insert_date,
data: old_remote_non_public_repeat_activity.data |> update_in(["to"], fn _ -> [] end)
})
|> Repo.update!()
assert length(Repo.all(Object)) == 3
Mix.Tasks.Pleroma.Database.run(["prune_objects", "--keep-threads", "--keep-non-public"])
Repo.all(Pleroma.Activity)
assert length(Repo.all(Object)) == 2
refute Object.get_by_ap_id(old_remote_post2_id)
end
+ test "with the --keep-threads option it keeps old threads a local user quoted", %{
+ old_insert_date: old_insert_date
+ } do
+ remote_user = insert(:user, local: false)
+ local_user = insert(:user)
+
+ {:ok, quoted} = CommonAPI.post(remote_user, %{status: "quote me"})
+
+ quoted
+ |> Ecto.Changeset.change(%{local: false, updated_at: old_insert_date})
+ |> Repo.update!()
+
+ {:ok, quote} = CommonAPI.post(local_user, %{status: "look", quote_id: quoted.id})
+
+ quote
+ |> Ecto.Changeset.change(%{updated_at: old_insert_date})
+ |> Repo.update!()
+
+ Mix.Tasks.Pleroma.Database.run(["prune_objects", "--keep-threads"])
+
+ assert Object.get_by_ap_id(quoted.data["object"])
+ end
+
test "with the --keep-threads option it keeps old threads addressed to a local user", %{
old_insert_date: old_insert_date
} do
remote_user = insert(:user, local: false)
local_user = insert(:user)
{:ok, mention} = CommonAPI.post(remote_user, %{status: "hey @#{local_user.nickname}"})
{:ok, dm} =
CommonAPI.post(remote_user, %{
status: "psst @#{local_user.nickname}",
visibility: "direct"
})
# The pin must not depend on the notifications still being there
Pleroma.Notification.clear(local_user)
for activity <- [mention, dm] do
activity
|> Ecto.Changeset.change(%{local: false, updated_at: old_insert_date})
|> Repo.update!()
end
Mix.Tasks.Pleroma.Database.run(["prune_objects", "--keep-threads"])
assert Object.get_by_ap_id(mention.data["object"])
assert Object.get_by_ap_id(dm.data["object"])
end
test "with the --keep-threads option it still keeps non-old threads even with no local interactions" do
remote_user = insert(:user, local: false)
remote_user2 = insert(:user, local: false)
{:ok, remote_post_activity} =
CommonAPI.post(remote_user, %{status: "some thing", local: false})
{:ok, remote_post_reply_activity} =
CommonAPI.post(remote_user2, %{
status: "some reply",
in_reply_to_status_id: remote_post_activity.id
})
remote_post_activity
|> Ecto.Changeset.change(%{local: false})
|> Repo.update!()
remote_post_reply_activity
|> Ecto.Changeset.change(%{local: false})
|> Repo.update!()
assert length(Repo.all(Object)) == 2
Mix.Tasks.Pleroma.Database.run(["prune_objects", "--keep-threads"])
assert length(Repo.all(Object)) == 2
end
test "with the --keep-threads option it deletes old threads with no local interaction", %{
old_insert_date: old_insert_date
} do
remote_user = insert(:user, local: false)
remote_user2 = insert(:user, local: false)
{:ok, old_remote_post_activity} =
CommonAPI.post(remote_user, %{status: "some thing", local: false})
old_remote_post_activity
|> Ecto.Changeset.change(%{local: false, updated_at: old_insert_date})
|> Repo.update!()
{:ok, old_remote_post_reply_activity} =
CommonAPI.post(remote_user2, %{
status: "some reply",
in_reply_to_status_id: old_remote_post_activity.id
})
old_remote_post_reply_activity
|> Ecto.Changeset.change(%{local: false, updated_at: old_insert_date})
|> Repo.update!()
{:ok, old_favourite_activity} =
CommonAPI.favorite(old_remote_post_activity.id, remote_user2)
old_favourite_activity
|> Ecto.Changeset.change(%{local: false, updated_at: old_insert_date})
|> Repo.update!()
{:ok, old_repeat_activity} = CommonAPI.repeat(old_remote_post_activity.id, remote_user2)
old_repeat_activity
|> Ecto.Changeset.change(%{local: false, updated_at: old_insert_date})
|> Repo.update!()
assert length(Repo.all(Object)) == 2
Mix.Tasks.Pleroma.Database.run(["prune_objects", "--keep-threads"])
assert length(Repo.all(Object)) == 0
end
test "with the --keep-threads option it keeps old threads with local interaction", %{
old_insert_date: old_insert_date
} do
remote_user = insert(:user, local: false)
local_user = insert(:user, local: true)
# local reply
{:ok, old_remote_post1_activity} =
CommonAPI.post(remote_user, %{status: "some thing", local: false})
old_remote_post1_activity
|> Ecto.Changeset.change(%{local: false, updated_at: old_insert_date})
|> Repo.update!()
{:ok, old_local_post2_reply_activity} =
CommonAPI.post(local_user, %{
status: "some reply",
in_reply_to_status_id: old_remote_post1_activity.id
})
old_local_post2_reply_activity
|> Ecto.Changeset.change(%{local: true, updated_at: old_insert_date})
|> Repo.update!()
# local Like
{:ok, old_remote_post3_activity} =
CommonAPI.post(remote_user, %{status: "some thing", local: false})
old_remote_post3_activity
|> Ecto.Changeset.change(%{local: false, updated_at: old_insert_date})
|> Repo.update!()
{:ok, old_favourite_activity} = CommonAPI.favorite(old_remote_post3_activity.id, local_user)
old_favourite_activity
|> Ecto.Changeset.change(%{local: true, updated_at: old_insert_date})
|> Repo.update!()
# local Announce
{:ok, old_remote_post4_activity} =
CommonAPI.post(remote_user, %{status: "some thing", local: false})
old_remote_post4_activity
|> Ecto.Changeset.change(%{local: false, updated_at: old_insert_date})
|> Repo.update!()
{:ok, old_repeat_activity} = CommonAPI.repeat(old_remote_post4_activity.id, local_user)
old_repeat_activity
|> Ecto.Changeset.change(%{local: true, updated_at: old_insert_date})
|> Repo.update!()
assert length(Repo.all(Object)) == 4
Mix.Tasks.Pleroma.Database.run(["prune_objects", "--keep-threads"])
assert length(Repo.all(Object)) == 4
end
test "with the --keep-threads option it keeps old threads with bookmarked posts", %{
old_insert_date: old_insert_date
} do
remote_user = insert(:user, local: false)
local_user = insert(:user, local: true)
{:ok, old_remote_post_activity} =
CommonAPI.post(remote_user, %{status: "some thing", local: false})
old_remote_post_activity
|> Ecto.Changeset.change(%{local: false, updated_at: old_insert_date})
|> Repo.update!()
Pleroma.Bookmark.create(local_user.id, old_remote_post_activity.id)
assert length(Repo.all(Object)) == 1
Mix.Tasks.Pleroma.Database.run(["prune_objects", "--keep-threads"])
assert length(Repo.all(Object)) == 1
end
test "We don't have unexpected tables which may contain objects that are referenced by activities" do
# We can delete orphaned activities. For that we look for the objects
# they reference in the 'objects', 'activities', and 'users' table.
# If someone adds another table with objects (idk, maybe with separate
# relations, or collections or w/e), then we need to make sure we
# add logic for that in the 'prune_objects' task so that we don't
# wrongly delete their corresponding activities.
# So when someone adds (or removes) a table, this test will fail.
# Either the table contains objects which can be referenced from the
# activities table
# => in that case the prune_objects job should be adapted so we don't
# delete activities who still have the referenced object.
# Or it doesn't contain objects which can be referenced from the activities table
# => in that case you can add/remove the table to/from this (sorted) list.
assert Repo.query!(
"SELECT table_name FROM information_schema.tables WHERE table_schema='public' AND table_type='BASE TABLE';"
).rows
|> Enum.sort() == [
["activities"],
["announcement_read_relationships"],
["announcements"],
["apps"],
["backups"],
["bookmark_folders"],
["bookmarks"],
["chat_message_references"],
["chats"],
["config"],
["conversation_participation_recipient_ships"],
["conversation_participations"],
["conversations"],
["counter_cache"],
["data_migration_failed_ids"],
["data_migrations"],
["deliveries"],
["filters"],
["following_relationships"],
["hashtags"],
["hashtags_objects"],
["instances"],
["lists"],
["markers"],
["mfa_tokens"],
["moderation_log"],
["notifications"],
["oauth_authorizations"],
["oauth_tokens"],
["oban_jobs"],
["oban_peers"],
["objects"],
["password_reset_tokens"],
["push_subscriptions"],
["registrations"],
["report_notes"],
["retention_cursors"],
["rich_media_card"],
["rules"],
["scheduled_activities"],
["schema_migrations"],
["thread_mutes"],
["user_follows_hashtag"],
# ["user_frontend_setting_profiles"], # not in pleroma
["user_invite_tokens"],
["user_notes"],
["user_relationships"],
["users"],
["webhooks"]
]
end
test "it prunes orphaned activities with the --prune-orphaned-activities" do
# Add a remote activity which references an Object
%Object{} |> Map.merge(%{data: %{"id" => "object_for_activity"}}) |> Repo.insert()
%Activity{}
|> Map.merge(%{
local: false,
data: %{"id" => "remote_activity_with_object", "object" => "object_for_activity"}
})
|> Repo.insert()
# Add a remote activity which references an activity
%Activity{}
|> Map.merge(%{
local: false,
data: %{
"id" => "remote_activity_with_activity",
"object" => "remote_activity_with_object"
}
})
|> Repo.insert()
# Add a remote activity which references an Actor
%User{} |> Map.merge(%{ap_id: "actor"}) |> Repo.insert()
%Activity{}
|> Map.merge(%{
local: false,
data: %{"id" => "remote_activity_with_actor", "object" => "actor"}
})
|> Repo.insert()
# Add a remote activity without existing referenced object, activity or actor
%Activity{}
|> Map.merge(%{
local: false,
data: %{
"id" => "remote_activity_without_existing_referenced_object",
"object" => "non_existing"
}
})
|> Repo.insert()
# Add a local activity without existing referenced object, activity or actor
%Activity{}
|> Map.merge(%{
local: true,
data: %{"id" => "local_activity_with_actor", "object" => "non_existing"}
})
|> Repo.insert()
# The remote activities without existing reference,
# and only the remote activities without existing reference, are deleted
# if, and only if, we provide the --prune-orphaned-activities option
assert length(Repo.all(Activity)) == 5
Mix.Tasks.Pleroma.Database.run(["prune_objects"])
assert length(Repo.all(Activity)) == 5
Mix.Tasks.Pleroma.Database.run(["prune_objects", "--prune-orphaned-activities"])
activities = Repo.all(Activity)
assert "remote_activity_without_existing_referenced_object" not in Enum.map(
activities,
fn a -> a.data["id"] end
)
assert length(activities) == 4
end
test "it prunes orphaned activities with the --prune-orphaned-activities when the objects are referenced from an array" do
%Object{} |> Map.merge(%{data: %{"id" => "existing_object"}}) |> Repo.insert()
%User{} |> Map.merge(%{ap_id: "existing_actor"}) |> Repo.insert()
# Multiple objects, one object exists (keep)
%Activity{}
|> Map.merge(%{
local: false,
data: %{
"id" => "remote_activity_existing_object",
"object" => ["non_ existing_object", "existing_object"]
}
})
|> Repo.insert()
# Multiple objects, one actor exists (keep)
%Activity{}
|> Map.merge(%{
local: false,
data: %{
"id" => "remote_activity_existing_actor",
"object" => ["non_ existing_object", "existing_actor"]
}
})
|> Repo.insert()
# Multiple objects, one activity exists (keep)
%Activity{}
|> Map.merge(%{
local: false,
data: %{
"id" => "remote_activity_existing_activity",
"object" => ["non_ existing_object", "remote_activity_existing_actor"]
}
})
|> Repo.insert()
# Multiple objects none exist (prune)
%Activity{}
|> Map.merge(%{
local: false,
data: %{
"id" => "remote_activity_without_existing_referenced_object",
"object" => ["owo", "whats_this"]
}
})
|> Repo.insert()
assert length(Repo.all(Activity)) == 4
Mix.Tasks.Pleroma.Database.run(["prune_objects"])
assert length(Repo.all(Activity)) == 4
Mix.Tasks.Pleroma.Database.run(["prune_objects", "--prune-orphaned-activities"])
activities = Repo.all(Activity)
assert length(activities) == 3
assert "remote_activity_without_existing_referenced_object" not in Enum.map(
activities,
fn a -> a.data["id"] end
)
assert length(activities) == 3
end
test "it prunes hashtags with no objects associated", %{old_insert_date: old_insert_date} do
user = insert(:user)
{:ok, hashtag_post_activity} =
CommonAPI.post(user, %{status: "morning #cofe", local: true})
hashtag_post_object = Object.normalize(hashtag_post_activity)
{:ok, hashtag_post2_activity} =
CommonAPI.post(user, %{status: "morning #cawfee", local: true})
hashtag_post2_object = Object.normalize(hashtag_post2_activity)
hashtag_post_object
|> Ecto.Changeset.change(%{updated_at: old_insert_date})
|> Repo.update!()
hashtag_post2_object
|> Ecto.Changeset.change(%{updated_at: old_insert_date})
|> Repo.update!()
# Test whether hashtags with follow relationships are kept
User.follow_hashtag(user, Hashtag.get_by_name("cofe"))
assert length(Repo.all(Hashtag)) == 2
assert length(Repo.all(Object)) == 2
Mix.Tasks.Pleroma.Database.run(["prune_objects"])
assert length(Repo.all(Hashtag)) == 1
assert length(Repo.all(Object)) == 0
assert Repo.one(Hashtag) |> Map.fetch!(:name) == "cofe"
end
end
describe "running update_users_following_followers_counts" do
test "following and followers count are updated" do
[user, user2] = insert_pair(:user)
{:ok, %User{} = user, _user2} = User.follow(user, user2)
following = User.following(user)
assert length(following) == 2
assert user.follower_count == 0
{:ok, user} =
user
|> Ecto.Changeset.change(%{follower_count: 3})
|> Repo.update()
assert user.follower_count == 3
assert {:ok, :ok} ==
Mix.Tasks.Pleroma.Database.run(["update_users_following_followers_counts"])
user = User.get_by_id(user.id)
assert length(User.following(user)) == 2
assert user.follower_count == 0
end
end
describe "running fix_likes_collections" do
test "it turns OrderedCollection likes into empty arrays" do
[user, user2] = insert_pair(:user)
{:ok, %{id: id, object: object}} = CommonAPI.post(user, %{status: "test"})
{:ok, %{object: object2}} = CommonAPI.post(user, %{status: "test test"})
CommonAPI.favorite(id, user2)
likes = %{
"first" =>
"http://mastodon.example.org/objects/dbdbc507-52c8-490d-9b7c-1e1d52e5c132/likes?page=1",
"id" => "http://mastodon.example.org/objects/dbdbc507-52c8-490d-9b7c-1e1d52e5c132/likes",
"totalItems" => 3,
"type" => "OrderedCollection"
}
new_data = Map.put(object2.data, "likes", likes)
object2
|> Ecto.Changeset.change(%{data: new_data})
|> Repo.update()
assert length(Object.get_by_id(object.id).data["likes"]) == 1
assert is_map(Object.get_by_id(object2.id).data["likes"])
assert :ok == Mix.Tasks.Pleroma.Database.run(["fix_likes_collections"])
assert length(Object.get_by_id(object.id).data["likes"]) == 1
assert Enum.empty?(Object.get_by_id(object2.id).data["likes"])
end
end
describe "ensure_expiration" do
test "it adds to expiration old statuses" do
activity1 = insert(:note_activity)
{:ok, inserted_at, 0} = DateTime.from_iso8601("2015-01-23T23:50:07Z")
activity2 = insert(:note_activity, %{inserted_at: inserted_at})
%{id: activity_id3} = insert(:note_activity)
expires_at = DateTime.add(DateTime.utc_now(), 60 * 61)
Pleroma.Workers.PurgeExpiredActivity.enqueue(
%{
activity_id: activity_id3
},
scheduled_at: expires_at
)
Mix.Tasks.Pleroma.Database.run(["ensure_expiration"])
assert_enqueued(
worker: Pleroma.Workers.PurgeExpiredActivity,
args: %{activity_id: activity1.id},
scheduled_at:
activity1.inserted_at
|> DateTime.from_naive!("Etc/UTC")
|> Timex.shift(days: 365)
)
assert_enqueued(
worker: Pleroma.Workers.PurgeExpiredActivity,
args: %{activity_id: activity2.id},
scheduled_at:
activity2.inserted_at
|> DateTime.from_naive!("Etc/UTC")
|> Timex.shift(days: 365)
)
assert_enqueued(
worker: Pleroma.Workers.PurgeExpiredActivity,
args: %{activity_id: activity_id3},
scheduled_at: expires_at
)
end
end
end
diff --git a/test/pleroma/retention_test.exs b/test/pleroma/retention_test.exs
index e6f037d4c..3bd7e115a 100644
--- a/test/pleroma/retention_test.exs
+++ b/test/pleroma/retention_test.exs
@@ -1,572 +1,587 @@
# Pleroma: A lightweight social networking server
# Copyright © 2017-2022 Pleroma Authors <https://pleroma.social/>
# SPDX-License-Identifier: AGPL-3.0-only
defmodule Pleroma.RetentionTest do
use Pleroma.DataCase, async: false
alias Pleroma.Activity
alias Pleroma.Bookmark
alias Pleroma.Object
alias Pleroma.Repo
alias Pleroma.Retention
alias Pleroma.Web.CommonAPI
import Bitwise
import Ecto.Query
import Pleroma.Factory
setup do
clear_config([:instance, :remote_post_retention_days], 90)
clear_config([:retention, :enabled], true)
clear_config([:retention, :max_objects], nil)
clear_config([:retention, :batch_size], 500)
clear_config([:retention, :keep_non_public], false)
old_date =
NaiveDateTime.utc_now()
|> NaiveDateTime.add(-200 * 86_400)
|> NaiveDateTime.truncate(:second)
%{old_date: old_date}
end
# Activity ids are time-ordered flakes and the retention walk goes by id, so
# an "old" activity needs an old id as well as an old timestamp. Must be
# applied before anything references the activity (bookmarks, notifications).
defp age(%Activity{} = activity, old_date, changes) do
ms = old_date |> DateTime.from_naive!("Etc/UTC") |> DateTime.to_unix(:millisecond)
old_id =
FlakeId.to_string(<<ms::integer-size(64), :rand.uniform(1 <<< 62)::integer-size(64)>>)
changes = Map.merge(%{id: old_id, updated_at: old_date}, changes)
{1, _} =
Activity
|> where([a], a.id == ^activity.id)
|> Repo.update_all(set: Map.to_list(changes))
Map.merge(activity, changes)
rescue
e in Postgrex.Error ->
reraise "age/3 must run before anything references the activity: #{Exception.message(e)}",
__STACKTRACE__
end
# Posts made through CommonAPI are local; turn the activity into a remote one
# and backdate it so it looks like an old federated post.
defp make_remote_and_old(activity, old_date), do: age(activity, old_date, %{local: false})
defp days_ago(days) do
NaiveDateTime.utc_now()
|> NaiveDateTime.add(-days * 86_400)
|> NaiveDateTime.truncate(:second)
end
defp make_old(activity, old_date), do: age(activity, old_date, %{})
defp old_remote_post(user, old_date, params \\ %{}) do
{:ok, activity} = CommonAPI.post(user, Map.merge(%{status: "some thing"}, params))
make_remote_and_old(activity, old_date)
end
describe "estimated_object_count/0" do
test "falls back to an exact count when the planner has no statistics" do
insert(:note)
insert(:note)
assert Retention.estimated_object_count() == Repo.aggregate(Object, :count)
end
end
describe "run/1" do
test "evicts an old remote thread nobody local touched", %{old_date: old_date} do
remote_user = insert(:user, local: false)
remote_user2 = insert(:user, local: false)
post = old_remote_post(remote_user, old_date)
{:ok, reply} =
CommonAPI.post(remote_user2, %{status: "reply", in_reply_to_status_id: post.id})
make_remote_and_old(reply, old_date)
{:ok, like} = CommonAPI.favorite(post.id, remote_user2)
make_remote_and_old(like, old_date)
assert %{contexts: 1, objects: 2, activities: 3} = Retention.run()
refute Object.get_by_ap_id(post.data["object"])
refute Object.get_by_ap_id(reply.data["object"])
assert Repo.all(Activity) == []
end
test "keeps recent remote threads" do
remote_user = insert(:user, local: false)
{:ok, post} = CommonAPI.post(remote_user, %{status: "fresh"})
post |> Ecto.Changeset.change(%{local: false}) |> Repo.update!()
assert %{contexts: 0, objects: 0} = Retention.run()
assert Object.get_by_ap_id(post.data["object"])
end
test "keeps old local posts", %{old_date: old_date} do
user = insert(:user)
{:ok, post} = CommonAPI.post(user, %{status: "mine"})
make_old(post, old_date)
assert %{contexts: 0, objects: 0} = Retention.run()
assert Object.get_by_ap_id(post.data["object"])
end
test "keeps an old remote thread a local user favourited", %{old_date: old_date} do
remote_user = insert(:user, local: false)
local_user = insert(:user)
post = old_remote_post(remote_user, old_date)
{:ok, like} = CommonAPI.favorite(post.id, local_user)
make_old(like, old_date)
assert %{contexts: 0} = Retention.run()
assert Object.get_by_ap_id(post.data["object"])
end
test "keeps an old remote thread a local user replied to", %{old_date: old_date} do
remote_user = insert(:user, local: false)
local_user = insert(:user)
post = old_remote_post(remote_user, old_date)
{:ok, reply} =
CommonAPI.post(local_user, %{status: "reply", in_reply_to_status_id: post.id})
make_old(reply, old_date)
assert %{contexts: 0} = Retention.run()
assert Object.get_by_ap_id(post.data["object"])
end
test "keeps an old remote thread a local user bookmarked", %{old_date: old_date} do
remote_user = insert(:user, local: false)
local_user = insert(:user)
post = old_remote_post(remote_user, old_date)
{:ok, _} = Bookmark.create(local_user.id, post.id)
assert %{contexts: 0} = Retention.run()
assert Object.get_by_ap_id(post.data["object"])
end
test "keeps an old remote post that mentioned a local user", %{old_date: old_date} do
remote_user = insert(:user, local: false)
local_user = insert(:user)
{:ok, post} = CommonAPI.post(remote_user, %{status: "hey @#{local_user.nickname}"})
Repo.delete_all(Pleroma.Notification)
post = make_remote_and_old(post, old_date)
assert {:ok, [_]} = Pleroma.Notification.create_notifications(post)
assert %{contexts: 0} = Retention.run()
assert Object.get_by_ap_id(post.data["object"])
end
test "keeps old remote posts addressed to a local user after notifications are cleared", %{
old_date: old_date
} do
remote_user = insert(:user, local: false)
local_user = insert(:user)
{:ok, dm} =
CommonAPI.post(remote_user, %{
status: "psst @#{local_user.nickname}",
visibility: "direct"
})
{:ok, mention} = CommonAPI.post(remote_user, %{status: "hey @#{local_user.nickname}"})
Pleroma.Notification.clear(local_user)
for post <- [dm, mention] do
make_remote_and_old(post, old_date)
end
assert %{contexts: 0} = Retention.run()
assert Object.get_by_ap_id(dm.data["object"])
assert Object.get_by_ap_id(mention.data["object"])
end
test "keeps a pinned post that an evicted thread's activity points at", %{
old_date: old_date
} do
remote_user = insert(:user, local: false)
remote_user2 = insert(:user, local: false)
local_user = insert(:user)
post = old_remote_post(remote_user, old_date)
{:ok, local_like} = CommonAPI.favorite(post.id, local_user)
make_old(local_like, old_date)
# A remote like that arrived with a context of its own
{:ok, like} = CommonAPI.favorite(post.id, remote_user2)
other_context = "https://remote.example/contexts/#{Ecto.UUID.generate()}"
age(like, old_date, %{local: false, data: Map.put(like.data, "context", other_context)})
assert %{contexts: 1, objects: 0} = Retention.run()
assert Object.get_by_ap_id(post.data["object"])
assert Activity.get_by_id(post.id)
end
test "keeps threads pinned by a local reaction of the legacy EmojiReaction type", %{
old_date: old_date
} do
remote_user = insert(:user, local: false)
local_user = insert(:user)
post = old_remote_post(remote_user, old_date)
Repo.insert!(%Activity{
data: %{
"id" => "#{Pleroma.Web.Endpoint.url()}/activities/#{Ecto.UUID.generate()}",
"type" => "EmojiReaction",
"actor" => local_user.ap_id,
"object" => post.data["object"],
"content" => "👍",
"context" => post.data["context"]
},
local: true,
actor: local_user.ap_id,
recipients: [remote_user.ap_id]
})
|> make_old(old_date)
assert %{contexts: 0} = Retention.run()
assert Object.get_by_ap_id(post.data["object"])
end
+ test "keeps an old remote thread whose post a local user quoted", %{old_date: old_date} do
+ remote_user = insert(:user, local: false)
+ local_user = insert(:user)
+
+ quoted = old_remote_post(remote_user, old_date)
+ {:ok, quote} = CommonAPI.post(local_user, %{status: "look at this", quote_id: quoted.id})
+ make_old(quote, old_date)
+
+ # The quote has a context of its own and does not pin the quoted thread.
+ refute quote.data["context"] == quoted.data["context"]
+
+ assert %{contexts: 0} = Retention.run()
+ assert Object.get_by_ap_id(quoted.data["object"])
+ end
+
test "keeps an old remote post that was reported", %{old_date: old_date} do
remote_user = insert(:user, local: false)
local_user = insert(:user)
post = old_remote_post(remote_user, old_date)
{:ok, _flag} =
CommonAPI.report(local_user, %{account_id: remote_user.id, status_ids: [post.id]})
assert %{contexts: 0} = Retention.run()
assert Object.get_by_ap_id(post.data["object"])
end
test "evicts an old remote thread even if remote users interacted with it", %{
old_date: old_date
} do
remote_user = insert(:user, local: false)
remote_user2 = insert(:user, local: false)
post = old_remote_post(remote_user, old_date)
{:ok, announce} = CommonAPI.repeat(post.id, remote_user2)
make_remote_and_old(announce, old_date)
assert %{contexts: 1, objects: 1, activities: 2} = Retention.run()
refute Object.get_by_ap_id(post.data["object"])
end
test "evicts old threads in batches, oldest first", %{old_date: old_date} do
clear_config([:retention, :batch_size], 1)
remote_user = insert(:user, local: false)
older = old_remote_post(remote_user, NaiveDateTime.add(old_date, -86_400))
newer = old_remote_post(remote_user, old_date)
assert %{contexts: 1} = Retention.run()
refute Object.get_by_ap_id(older.data["object"])
assert Object.get_by_ap_id(newer.data["object"])
assert %{contexts: 1} = Retention.run()
refute Object.get_by_ap_id(newer.data["object"])
end
test "evicts threads younger than the retention period when over the object watermark" do
clear_config([:retention, :max_objects], 1)
clear_config([:retention, :batch_size], 1)
remote_user = insert(:user, local: false)
local_user = insert(:user)
older = old_remote_post(remote_user, days_ago(3))
newer = old_remote_post(remote_user, days_ago(2))
{:ok, mine} = CommonAPI.post(local_user, %{status: "mine"})
assert %{contexts: 1} = Retention.run()
refute Object.get_by_ap_id(older.data["object"])
assert Object.get_by_ap_id(newer.data["object"])
assert Object.get_by_ap_id(mine.data["object"])
end
test "keeps threads quiet for less than a day, even over the watermark" do
clear_config([:retention, :max_objects], 0)
remote_user = insert(:user, local: false)
post = old_remote_post(remote_user, NaiveDateTime.add(days_ago(1), 3_600))
assert %{contexts: 0} = Retention.run()
assert Object.get_by_ap_id(post.data["object"])
end
test "keeps an old thread with a recent reply, even over the watermark", %{
old_date: old_date
} do
clear_config([:retention, :max_objects], 0)
remote_user = insert(:user, local: false)
{:ok, post} = CommonAPI.post(remote_user, %{status: "old but active"})
# Old id, so the walk reaches it, but active an hour ago
age(post, old_date, %{local: false, updated_at: NaiveDateTime.add(days_ago(0), -3_600)})
assert %{contexts: 0} = Retention.run()
assert Object.get_by_ap_id(post.data["object"])
end
test "does not evict anything below the watermark" do
clear_config([:retention, :max_objects], 100)
remote_user = insert(:user, local: false)
{:ok, post} = CommonAPI.post(remote_user, %{status: "fresh"})
post |> Ecto.Changeset.change(%{local: false}) |> Repo.update!()
assert %{contexts: 0} = Retention.run()
assert Object.get_by_ap_id(post.data["object"])
end
test "with keep_non_public it keeps old non-public remote threads", %{old_date: old_date} do
clear_config([:retention, :keep_non_public], true)
remote_user = insert(:user, local: false)
{:ok, private} = CommonAPI.post(remote_user, %{status: "psst", visibility: "private"})
make_remote_and_old(private, old_date)
public = old_remote_post(remote_user, old_date)
assert %{contexts: 1} = Retention.run()
assert Object.get_by_ap_id(private.data["object"])
refute Object.get_by_ap_id(public.data["object"])
end
test "never touches activities without a context, even over the watermark" do
clear_config([:retention, :max_objects], 1)
remote_user = insert(:user, local: false)
remote_user2 = insert(:user, local: false)
local_user = insert(:user)
{:ok, _, _, follow} = CommonAPI.follow(local_user, remote_user)
{:ok, block} = CommonAPI.block(remote_user2, local_user)
# Enough objects to trip the watermark
{:ok, post} = CommonAPI.post(remote_user, %{status: "a"})
post |> Ecto.Changeset.change(%{local: false}) |> Repo.update!()
{:ok, post2} = CommonAPI.post(remote_user, %{status: "b"})
post2 |> Ecto.Changeset.change(%{local: false}) |> Repo.update!()
for activity <- [follow, block] do
assert is_nil(activity.data["context"])
activity
|> Ecto.Changeset.change(%{local: false})
|> Repo.update!()
end
Retention.run()
assert Activity.get_by_id(follow.id)
assert Activity.get_by_id(block.id)
end
test "keeps remote reports and the threads they point at", %{old_date: old_date} do
remote_user = insert(:user, local: false)
reporter = insert(:user, local: false)
admin = insert(:user, is_admin: true)
post = old_remote_post(remote_user, old_date)
flag =
Repo.insert!(%Activity{
data: %{
"id" => "https://domain1.com/activities/#{Ecto.UUID.generate()}",
"type" => "Flag",
"actor" => reporter.ap_id,
"object" => [remote_user.ap_id, post.data["object"]],
"context" => "https://domain1.com/contexts/#{Ecto.UUID.generate()}",
"content" => "spam"
},
local: false,
actor: reporter.ap_id,
recipients: []
})
flag = make_old(flag, old_date)
{:ok, _} = Pleroma.ReportNote.create(admin.id, flag.id, "looking into it")
assert %{contexts: 0} = Retention.run()
assert Activity.get_by_id(flag.id)
assert Object.get_by_ap_id(post.data["object"])
end
test "removes remote activities outside the thread that point at evicted objects", %{
old_date: old_date
} do
remote_user = insert(:user, local: false)
post = old_remote_post(remote_user, old_date)
delete =
Repo.insert!(%Activity{
data: %{
"id" => "https://domain1.com/activities/#{Ecto.UUID.generate()}",
"type" => "Delete",
"actor" => remote_user.ap_id,
"object" => post.data["object"]
},
local: false,
actor: remote_user.ap_id,
recipients: []
})
assert %{contexts: 1, objects: 1, activities: 2} = Retention.run()
refute Activity.get_by_id(delete.id)
end
test "never deletes local objects, even inside an otherwise remote thread", %{
old_date: old_date
} do
local_user = insert(:user)
{:ok, post} = CommonAPI.post(local_user, %{status: "mine, honestly"})
# Pretend the Create activity federated in from elsewhere, addressed to
# nobody local (the local followers collection would pin the thread)
age(post, old_date, %{local: false, recipients: []})
assert %{contexts: 1, objects: 0} = Retention.run()
assert Object.get_by_ap_id(post.data["object"])
end
test "remembers where it got to and does not rescan", %{old_date: old_date} do
clear_config([:retention, :batch_size], 1)
remote_user = insert(:user, local: false)
post = old_remote_post(remote_user, old_date)
assert is_nil(Retention.Cursor.get("activities"))
assert %{contexts: 1} = Retention.run()
assert Retention.Cursor.get("activities") == post.id
# An older activity that appears behind the cursor is not visited again
# until the cursor is reset.
older = old_remote_post(remote_user, NaiveDateTime.add(old_date, -86_400))
assert %{contexts: 0} = Retention.run()
assert Object.get_by_ap_id(older.data["object"])
Retention.Cursor.reset("activities")
assert %{contexts: 1} = Retention.run()
refute Object.get_by_ap_id(older.data["object"])
end
test "keeps an old thread that is still active", %{old_date: old_date} do
remote_user = insert(:user, local: false)
remote_user2 = insert(:user, local: false)
post = old_remote_post(remote_user, old_date)
{:ok, reply} =
CommonAPI.post(remote_user2, %{status: "still here", in_reply_to_status_id: post.id})
reply |> Ecto.Changeset.change(%{local: false}) |> Repo.update!()
assert %{contexts: 0} = Retention.run()
assert Object.get_by_ap_id(post.data["object"])
end
test "keeps threads pinned by each local interaction type", %{old_date: old_date} do
remote_user = insert(:user, local: false)
local_user = insert(:user)
liked = old_remote_post(remote_user, old_date)
{:ok, like} = CommonAPI.favorite(liked.id, local_user)
make_old(like, old_date)
repeated = old_remote_post(remote_user, old_date)
{:ok, announce} = CommonAPI.repeat(repeated.id, local_user)
make_old(announce, old_date)
reacted = old_remote_post(remote_user, old_date)
{:ok, react} = CommonAPI.react_with_emoji(reacted.id, local_user, "👍")
make_old(react, old_date)
replied = old_remote_post(remote_user, old_date)
{:ok, reply} = CommonAPI.post(local_user, %{status: "r", in_reply_to_status_id: replied.id})
make_old(reply, old_date)
assert %{contexts: 0} = Retention.run()
for post <- [liked, repeated, reacted, replied] do
assert Object.get_by_ap_id(post.data["object"])
end
end
test "a watermark run does not stall the age-based walk", %{old_date: old_date} do
remote_user = insert(:user, local: false)
clear_config([:retention, :max_objects], 0)
recent = old_remote_post(remote_user, days_ago(2))
assert %{contexts: 1} = Retention.run()
refute Object.get_by_ap_id(recent.data["object"])
clear_config([:retention, :max_objects], nil)
old = old_remote_post(remote_user, old_date)
assert %{contexts: 1} = Retention.run()
refute Object.get_by_ap_id(old.data["object"])
end
test "moves the cursor to the deadline once everything old is done", %{old_date: old_date} do
remote_user = insert(:user, local: false)
old_remote_post(remote_user, old_date)
Retention.run()
<<cursor_ms::integer-size(64), _::integer-size(64)>> =
FlakeId.from_string(Retention.Cursor.get("activities"))
deadline_ms =
Retention.deadline()
|> DateTime.from_naive!("Etc/UTC")
|> DateTime.to_unix(:millisecond)
assert_in_delta cursor_ms, deadline_ms, 5_000
end
test "removes hashtags that no longer have objects", %{old_date: old_date} do
remote_user = insert(:user, local: false)
old_remote_post(remote_user, old_date, %{status: "#lonelytag"})
assert Pleroma.Hashtag.get_by_name("lonelytag")
Retention.run()
refute Pleroma.Hashtag.get_by_name("lonelytag")
end
test "only cleans up hashtags of the evicted objects", %{old_date: old_date} do
remote_user = insert(:user, local: false)
local_user = insert(:user)
old_remote_post(remote_user, old_date, %{status: "#gone #shared #followed"})
{:ok, _} = CommonAPI.post(local_user, %{status: "#shared"})
{:ok, followed} = Pleroma.Hashtag.get_or_create_by_name("followed")
{:ok, _} = Pleroma.User.follow_hashtag(local_user, followed)
{:ok, _} = Pleroma.Hashtag.get_or_create_by_name("unrelated")
assert %{contexts: 1} = Retention.run()
refute Pleroma.Hashtag.get_by_name("gone")
assert Pleroma.Hashtag.get_by_name("shared")
assert Pleroma.Hashtag.get_by_name("followed")
assert Pleroma.Hashtag.get_by_name("unrelated")
end
end
end

File Metadata

Mime Type
text/x-diff
Expires
Fri, Oct 9, 4:23 AM (1 d, 8 h)
Storage Engine
blob
Storage Format
Raw Data
Storage Handle
1784437
Default Alt Text
(130 KB)

Event Timeline